The Mathocalypse

Last night my 9-year-old son was taunting my wife, complexity theorist Dana Moshkovitz, as follows: “mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!”

While my son was being a brat, he also wasn’t wrong. Whether you’re thrilled, depressed, angry, or whatever else about it, yesterday was surely one of the biggest days in mathematical history. And yes, among the 372 huge results released yesterday by OpenAI, on the recommendation of its advisory group of Timothy Gowers, Edward Witten, and other distinguished mathematicians, was a proof of Subhash Khot’s Unique Games Conjecture (UGC), a statement that my wife has worked toward proving for the entire time I’ve known her. (The UGC implies that a whole slew of optimization problems really are NP-hard, even if you just want an approximation that’s slightly better than what you get from semidefinite programming relaxation, which is one of our main tools.)

Or at least, we’re pretty sure that it’s a proof! There’s a Lean certificate, as there are for some of the other 372 breakthrough results (not all of them). But it also appears that no human has understood just about any of these proofs yet; the race to do so has just started. If you want an on-the-ground sense of what that race is going to be like, here’s some of what Dana texted me last night:

It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results

Basically the paper is so horribly written that it’s impossible to read it without AI help

I asked Astra for reasonable completeness and soundness claims of the noise gadget and it gave them by combining claims from all over the paper

They also have direct optimal NP hardness of approximation proofs for the main applications of the UGC (Max Cut and all CSP) that bypass the UGC.

The UGC proof invents a completely new bizarre code with a noise test. It’s some crazy recursive construction.

It’s not the long code, not the short code – some alien craziness

I still think that there maybe is a proof that uses the half space code (which is natural)

The citations are often irrelevant and confusing

A possible future is a math world that’s heavenly if you have vision/creative ideas that AI could help check and implement.

And of course there’s a lot for us to learn from the aliens

If you’re wondering what emotions Dana is feeling—well, probably all of them! Even while a central career aspiration has fallen to a robot, there are at least two mitigating factors for her. First, she can feel vindicated that the UGC was true after all, something she never doubted even while many of her colleagues did! Second, all of us in math and theoretical computer science and mathematical physics, at least those who cared about solving crisply-stated problems, are now in the same boat.


Besides the Unique Games Conjecture, here’s a small sampling of the treasures from Aladdin’s cave that I’ll probably be paying the most attention to over the coming weeks:

Any of the above, alone, could easily have been “result of the year” in some area (and in some cases, like Unique Games and L=BPL, in all of CS theory). And there’s a lot that I’ve left out—feel free to share in the comments whatever is making your eyes bug out! There are equally astounding wonders in number theory, combinatorics, algebraic geometry, analysis, and pretty much every other area of math, most of which I’ll never understand, although I’ll note that it includes partial progress toward the Riemann hypothesis and the Hodge Conjecture and the Birch-Swinnerton-Dyer Conjecture (i.e., the majority of the remaining Millennium Problems).

We can take solace in what’s missing from the list. P ≠NP isn’t there, nor even P=BPP or NEXP⊄P/poly, and surely not for lack of trying. Apparently the greatest open problems of theoretical computer science are indeed pretty hard!


Oh, lest I forget: one day before the OpenAI dump, meaning Monday evening, Virginia Williams and Josh Alman posted an arXiv preprint that solves the 3SUM problem in O(n1.9992) time, and the All-Pairs Shortest Paths problem in O(n2.9995) time, refuting half-century-old conjectures that the correct answers were n2-o(1) and n3-o(1) respectively. In this case, it wasn’t an OpenAI model that supplied the crucial idea; it was an Anthropic one! But Anthropic then took a different approach from OpenAI: rather than post the undigested solutions to the world, it gave Virginia and Josh the opportunity to write and announce a digested version in exchange for compensation.

These have emerged as the two main models for communicating AI math breakthroughs, and they both have strengths and weaknesses. The “OpenAI model” sets up a crazy race among humans to digest and explain a messy AI proof (work that could easily be some combination of thankless, barely-credited, competitive, and unfun), while the “Anthropic model” puts a private company in the position of picking and choosing which human mathematicians get to be the emissaries of the AI. Dunno, what do you guys think?


For those who are wondering: apparently, the AI model that produced all these wonders was not bespoke contraption of 10,000 agents burning millions of dollars worth of compute, as was used for example to construct a finite-time blowup for the Navier-Stokes equations. Instead, it was simply the latest internal OpenAI model—one that might be released to paying ChatGPT customers within the next couple of months, depending on the recommendations of OpenAI’s safety board! (My 9-year-old son: “Oh they definitely shouldn’t release that. If it could solve all those math problems, it can’t possibly be safe.”) Apparently they used about 3 hours of GPT-Pro level compute on average per problem solved.

Also, if you were wondering: apparently they tried the model on about 8,000 problems. So, right now it “merely” solves ~5% of the longstanding open mathematical problems that it’s asked about, the problems that whole communities have spent years on, after a single 3-hour attempt on them.


I’ve been glad to see the CS theory community rising to the occasion. At the Simons Institute in Berkeley, here at UT Austin, and elsewhere, I’ve hearing stories of researchers rushing to pore over the manuscripts and make sense of them and explain them—because what else do we do? How else do we continue the craft to which we’ve devoted much of our lives?

If you want some sense of what things feel like now in math, imagine a hunter-gatherer who’s spent his entire life learning to survive deep in an unforgiving rainforest, then a giant resort hotel springs up right next to him with a helipad and heated pools and AirBnBs, and without missing a beat, the hunter-gatherer says: “alright fine, so now my new job is to run wilderness retreats for the tourists, or something.”

In Quanta magazine, Jordana Cepelewitz attempted a different metaphor:

It’s as if you were teleported to the peak of a tall mountain. Surrounded by fog, you have no idea where you are, or what’s around you. You do not know how your mountain connects to others, and you have no equipment to help you explore, no way to help someone else join you. If you had climbed the mountain yourself, you would have experienced how the human body adapts to altitude and changes in oxygen levels. You might have had to invent tools to navigate, to climb steep cliffs, or to make a shelter. You might have encountered a fellow explorer, gotten lost together in a hidden valley, and found a plant that could be turned into a life-saving medicine.

Instead you’re perched on the peak but in the dark, while the maker of the teleportation machine tells you that it can explore the wilderness better than any human.

For any one of these mountains, if we care enough, I feel optimistic that we can do as we always have: clear the fog and figure out the path, except now using the teleportation machine to help guide us. The bigger challenge will be to nurture a community that still cares about the heroic adventure of finding the paths up these mountains in the world with the machine. (Oh, and I think one place where the metaphor breaks is that we still do have each other, as much as we ever did before!)


Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.

So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.” Or maybe none of the 372 well-known open problems that were solved were real math problems, they were all just glorified contest puzzles and trivialities. (After all, there’s still no Riemann Hypothesis!) Or maybe the entire 4000-year-old discipline of mathematics needs to be jettisoned: turns out that it was all just puzzle-solving and trivialities; all that’s different is that now the triviality stands unmasked. In any case, what really matters is that the true inner sanctum of human creativity hasn’t been breached and probably never will be, and also, that Sam Altman and Dario Amodei are contemptible little nerds.

If you’re still a proponent of that doomed worldview, still aboard the sinking ship, I encourage you in the strongest possible terms to read yesterday’s other great contribution to AI discourse, besides the OpenAI Mathocalypse dump: namely, Scott Alexander’s open letter to Steven Pinker. I feel some responsibility for this, as the person who first introduced Steven Pinker to the existence of the rationalist community, and who also first introduced Steven Pinker and Scott Alexander to one another (they had both been fans of each other’s writing). And now Scott is challenging Steve to a literal duel, with guns!

For whatever it’s worth: Steve is a lifelong intellectual hero of mine, just as he is for Scott, and I also have to privilege of calling Steve my friend. But I found Scott’s post to be one of the most devastating rejoinders to anything that I’ve ever read. And I thought Scott’s conclusion was exactly right: when it comes to AI risk, Steve’s great challenge is now to accept and start using a more “Pinkerite” epistemology.


Last night, while I should’ve been poring over some of OpenAI’s hundreds of papers and/or writing this post, I decided to spend some time with my kids instead. They wanted a movie night, so I suggested something they’d never seen before (and that I hadn’t seen for decades), and that seemed chock-full of no-nonsense, practical guidance for the world in which they’re going to grow up: Terminator 2.

16 Responses to “The Mathocalypse”

  1. Sniffnoy Says:

    Hey remember how just a few months ago Hong Wang won the Fields Medal for her work on proving the Kakeya Conjecture in 3 dimensions? Well now OpenAI says they’ve done it in 4 dimensions, just like that…

  2. Scott Says:

    Additional commentary from Dana on the UGC proof:

      The Unique Games Conjecture is true, though it wasn’t really needed for the proof of its most famous corollaries (like optimal hardness of approximating Max-Cut and other constraint satisfaction problems). The published proof of the conjecture somehow makes the short code approach* work with a bunch of new ideas (it’ll take time to understand exactly how, the writeup is poorly written). I still wonder whether it’s possible to prove the conjecture via the half-space code.

      * Addendum: Upon further inspection, the proof constructs a completely new tree-based code that somehow has a noise test with properties similar to that of the short code.

  3. Manu Says:

    I can only speak for the logic part, but the results are astonishing. All these famous problems solved in the blink of an eye. Pretty depressing news, in my opinion. The only future i see now for human math is a slow death. Lets only hope humanity itself fares better…

  4. Olav Says:

    It’s surreal to be living through this. My condolences (so to speak) to Dana for having one of her own problems crushed by the robot. But it’s interesting to hear her initial reactions to the proof—I hope the AI labs manage to train the models to write papers in a way that’s more understandable to humans in the not-too-distant future. Scott, if you go to the trouble of studying in some depth any of the results you mention in your post, it would be very interesting to hear your reactions and insights as well.

  5. Dana Moshkovitz Says:

    I don’t think that the lonely teleport to the peak of the mountain is a good metaphor, not just because we’re not lonely. For each of those problems we had a community, frameworks and existing results that the AI clearly built on. The AI results represent paradigm shifts, they do teleport us much farther, but they don’t finish math.

  6. Andrei Says:

    >one of the biggest days in mathematical history

    one of the biggest days in mathematical history so far…

    Honestly, I expect a continuous tsunami to arrive when these internal models are made publicly available.

  7. Pace Nielsen Says:

    To answer your question: I prefer the “let them all dig through the alien proofs” approach over the “select a few winners” approach. (Plus, as LLMs improve, I expect the expositions to improve.)

  8. Jeremy Says:

    It seems you have an editing issue, youve left in two different bullet points referring to maximum matching. This goes to show that these so called “AI math” results are just exaggeration and hype (jk jk).

  9. Nicolas Rougerie Says:

    Doing mathematical physics myself. Several of the main problems I would have called central to my subfield are claimed as solved. Certainly an earthquake, so don’t get me wrong … but …

    My first look into the preprints is disappointing: unreadable. Don’t know Lean myself so I won’t check the formalized proofs. For these kind of questions we knew the yes/no answer already (physics … experiments …) the point is “what kind of new idea was needed for a full proof.”

    So … certainly this is a revolution. The beginning of one anyway, for now we need to extract some sense out of all this mess. Unless we decide to dedicate our energy to pause AI development while we still can. Or focus on a real problem, like global warming. As of yet undecided …

  10. DavidM Says:

    > Apparently they used about 3 hours of GPT-Pro level compute on average per problem solved.

    This is quite amazing – is there a public source for this, or is it [personal communication]?

  11. Daniel Guyton Says:

    At least you chose Terminator 2 where the Terminator is Good and even Sarah Connor admits he’d be a good father.

  12. Tristan Says:

    It never ceases to astonish me how readily people vaunt the supposed powers of the humble LLM, which is, after all, nothing more than a next-token predictor: a probability distribution over tokens, a rather grandly advertised autocomplete. You persist in crediting it with faculties it manifestly does not possess. It merely interpolates patterns in its training data. It is intrinsically incapable of original insight, of conceiving genuinely new ideas, or of constructing new mathematical objects. One might have thought this obvious to anyone with even a passing understanding of how feed-forward neural networks are trained.

    No doubt all this does rather resemble a “mathpocalypse” to mathematicians of your particular calibre. But perhaps the occasion calls for a little reflection. If a pattern-interpolating autocomplete can produce your theorems quite so readily, one is entitled to wonder how intellectually demanding the enterprise was to begin with. Perhaps what you took for creativity and insight was, on closer inspection, neither.

    For my own part, I shall shortly be embarking upon a PhD in algebraic geometry, having finished near the top of my cohort in the Mathematical Tripos at Cambridge. I mention this merely to establish that my confidence is not entirely uninformed. Algebraic geometry is, mercifully, a field in which genuine creativity remains indispensable: one must actually conceive and construct new mathematical objects, rather than rearrange familiar patterns and dignify the exercise with the word “discovery”. It is precisely this sort of work that lies beyond the capacities of a glorified autocomplete.

    I therefore regard the much-advertised “mathpocalypse” with a certain equanimity. It may, at last, separate the wheat from the chaff, making rather more apparent the distinction between those who genuinely think mathematically and those who have merely become proficient at reproducing its outward forms.

    But I mustn’t keep you. My Noetherian schemes await.

    Cheerio!

  13. Ivo Says:

    Scott, is this a turning point where mathematics gets off its well trodden path of logic and becomes more like an empirical science? Heck, people were arguing if math is invented or discovered, this settles it. Math is discovered – by AI!

    From now on forth, Mathematicians will become scientists, wielding AI as a telescope, pointed towards new and disturbing alien worlds in the Platonic realm.

  14. Matt Says:

    Hi Scott, I am unfortunately on your side now, despite the fact that I never wanted to. I’m still not sure about ASI escaping containment but certainly it is true now that humans can do truly catastrophic things to our society with these results. I would not be surprised if they have already broken cryptography. As a young PhD student, please let me know what can be done here. I do not want to be left at the mercy of doomsday cultists and best and mobsters at worst. There is absolutely no way I can see the outcome of this business being even close to a positive one. I am less concerned about the consequences to mathematical research (we will probably adapt, if slowly) but rather to what this actually means for broader society. What if they have better factoring? Better Haber-Bosch? A practical improvement on FFT? And what if the NSA has these things? God knows what we can do about it.

  15. Scott Says:

    DavidM #10: Oh sorry, someone in my research group told me that had been made public!

    I’m headed to the airport right now — as it happens, headed to Vancouver, to give a talk at UBC about the AI Mathocalypse! A talk that I’ll now need to edit on my way there, given all the developments since I last gave the talk last week. 🙂

    So, can anyone find a source for the 3-hour claim? If not, I’ll edit the post to remove it.

  16. Scott Says:

    Tristan #12: LOL! That was a delightful parody (at least, I assume it’s a parody) of what I’ve now heard dozens of times in earnest.

Leave a Reply

You can use rich HTML in comments! You can also use basic TeX, by enclosing it within $$ $$ for displayed equations or \( \) for inline equations.

Comment Policies:

After two decades of mostly-open comments, in July 2024 Shtetl-Optimized transitioned to the following policy:

All comments are treated, by default, as personal missives to me, Scott Aaronson---with no expectation either that they'll appear on the blog or that I'll reply to them.

At my leisure and discretion, and in consultation with the Shtetl-Optimized Committee of Guardians, I'll put on the blog a curated selection of comments that I judge to be particularly interesting or to move the topic forward, and I'll do my best to answer those. But it will be more like Letters to the Editor. Anyone who feels unjustly censored is welcome to the rest of the Internet.

To the many who've asked me for this over the years, you're welcome!