OpenAI just dropped a bombshell on the mathematics world. The company's latest internal model, Astra, solved 10 longstanding mathematical problems that have stumped researchers for decades—breakthroughs ranging from quantum game theory to higher-dimensional sphere packing. According to OpenAI's announcement, these aren't obscure puzzles mathematicians ignore. They're problems that leading researchers have spent careers trying to crack. And now the field is having what Fields Medal winner James Maynard calls an existential crisis, questioning what happens to mathematics when AI can solve in weeks what humans couldn't solve in years.
OpenAI just upended mathematics. A few weeks ago, the company published what it called "10 Advances in Mathematics and Theoretical Computer Science"—a collection of solutions to problems that have resisted human mathematicians for years, some for decades. The work came from Astra, an unreleased internal model that appears to represent a dramatic leap in AI's reasoning capabilities.
The reaction from the mathematical community has been immediate and conflicted. "Shell shock" is how Robert Hart, The Verge's AI reporter, describes it after interviewing nearly a dozen leading mathematicians. "It's happened so quickly that it has just taken a lot of people by surprise," he explains.
What makes this different from previous AI breakthroughs is the caliber of work. These aren't simplified test problems or areas mathematicians consider irrelevant. "If a researcher had done any one of these problems, they'd probably be set for an academic career," multiple researchers told Hart. If a human had solved all 10, many wouldn't believe it.
The problems span disparate fields—quantum game theory, sphere packing in dimensions higher than three, and areas of theoretical computer science that require sophisticated abstract reasoning. According to those who've reviewed the hundreds of pages of proofs OpenAI published, the work is legitimate. The company even formalized the proofs in Lean, a programming language that can verify mathematical rigor computationally.
But here's where it gets weird: AI models are still terrible at elementary arithmetic. They can't reliably count, struggle with days of the week, and as recently as 2024, famously couldn't count the Rs in "strawberry". Yet they're now solving problems that require deep mathematical intuition.
"To be good at math, you've got to be good at counting, or adding, or multiplying," Hart notes. "A lot of it is actually reasoning. If you look at academic math papers, a lot of the time you won't see numbers." That disconnect—terrible at arithmetic, excellent at abstract proof—reveals something fundamental about how these models work, and what they're capable of.
The achievements also raise uncomfortable questions about attribution. OpenAI's initial blog post claimed these were 10 problems with "no progress in the last 10 years." But the papers themselves acknowledge building on recent work by named researchers. The company quietly changed the blog post after researchers pointed out the discrepancy. It's the kind of sloppiness that would end a human academic's career, but seems to get overlooked when the scale of discovery is this large.
Then there's the repeatability question. How many problems did Astra attempt before succeeding at these 10? OpenAI won't say. The company has hired cohorts of senior mathematicians, raising questions about how much human guidance shaped these results. Without access to the model or transparency about the process, independent verification is impossible.
For mathematicians, the crisis isn't just about solved problems. It's about what solving problems means for the field. James Maynard, who won the Fields Medal—mathematics' highest honor—told Hart he's been "soul-searching." Johannes Schmitt, a researcher in Zurich, worries about a future where "math problems get 'mowed down' by AI, but we don't actually push the field forward because humans are taken out of the loop."
That fear cuts to the heart of mathematical research. "The most interesting discoveries in the field aren't that you've solved something, it's what evolves from that," Maynard explained. Solutions open new fields of research, create new tools, pose new questions. But if AI just ticks off existing problems without that generative spark, it could leave mathematics sterile—a field of answered questions with nowhere left to go.
The timing couldn't be worse for graduate students. "If the standard for a publishable paper in math is something that an AI cannot do, particularly when a PhD is typically four years, the challenge is you're not trying to come up with a problem that AI can't do now, it's an AI in four years' time," Maynard said. Students picking research topics today are gambling that AI won't solve their thesis problem halfway through their doctorate.
Colva Roney-Dougal at St. Andrews captured another dimension of the crisis: cost and access. Mathematics has been one of the cheapest scientific disciplines. "A lot of the time I don't bother getting a research grant," she said. "I don't need one. I just have a blackboard." OpenAI claims these 10 results cost around $2,000 in compute—but that's a generous figure that doesn't account for model development, and it's still prohibitive for a field that traditionally operates on minimal budgets.
Roney-Dougal also expressed frustration with how AI labs are approaching mathematics. "They're treating our discipline as an advertising playground," she told Hart. Mathematics makes for clean demos—no messy lab experiments, no ethical concerns about human subjects, just pure logic that's easy to verify and impressive to announce. Whether AI companies care about advancing mathematical knowledge, or just want flashy benchmarks, remains unclear.
Some mathematicians have signed the Leiden Declaration, pledging not to buy into AI hype. But as Hart points out, that pledge isn't stopping the hype, and AI labs aren't slowing down their use of academic disciplines as marketing vehicles.
There's also a democratization argument. AI could give people with mathematical intuition but no formal training access to high-level problem-solving. Hart heard stories of talented undergraduates producing graduate-level work with AI assistance—work they never could have done alone. But he also heard complaints about AI-generated papers flooding journals and pre-print servers, submitted by people who lack the skills to verify whether they've actually solved anything.
Gary Marcus, a prominent AI skeptic, warns against generalizing from success in one domain. "As we learned a decade ago from AI's shambolic and ultimately failed attempt to turn Jeopardy-winning Watson into a cancer-fighting machine, success in one domain does not guarantee success in all," Marcus wrote. But even skeptics acknowledge there's been "an undeniable trajectory in the last few years of a broadening capability increase," as Hart puts it.
Andras Juhasz, a professor at Oxford, offered a telling observation: He doesn't think current AI has "any geometric intuition whatsoever," which might explain limited progress in fields like topology. AI capabilities remain jagged—exceptional in some areas, incompetent in others. But the question is where those capabilities peak, and how quickly the gaps close.
For now, mathematicians are caught between shell shock and uncertainty. The field hasn't had time to react, to restructure, to figure out what comes next. "Even they don't really know what's happening," Hart said of excited researchers. "There was still this lingering uncertainty of like, well, where does this leave the field?"
The mathematics community faces an inflection point. OpenAI's Astra has proven AI can solve problems that define academic careers, but questions about methodology, attribution, cost, and the future of mathematical research remain unanswered. As competing labs rush to replicate or exceed these results, mathematicians will need more than the Leiden Declaration to navigate what one researcher called treating their discipline as "an advertising playground." The real test isn't whether AI can solve problems—it's whether those solutions open new fields of inquiry or simply close old ones, leaving mathematics with answers but nowhere to go.