Dispatches from the possibly last days of human relevance

As most readers have presumably heard by now, Paul Erdös’s Unit Distance Problem from 1946—one of the central open problems from the field of discrete geometry—has been solved by an internal OpenAI model. Erdös had conjectured that, given n points in the plane, at most n1+o(1) pairs of them could be unit distance apart. Using high-powered results from algebraic number theory, GPT refuted this, constructing a set with n1+ε unit-distance pairs, for ε ~ 10-38. Shortly afterward, Will Sawin, a human (!), improved GPT’s construction to get ~n1.014 pairs. There’s since been a claim to improve this further, to ~n1.034. Meanwhile, the best known upper bound remains n4/3, improving Erdös’s n3/2.

The entire process seems have been one-shot: my former student Lijie Chen simply gave GPT the problem, then GPT thought for a while and output a several-page argument that, on analysis by human experts, turned out to be correct. Of course there’s selection bias here; we’re not hearing as much about the hundreds of other problems GPT was given that it didn’t solve (isn’t that the case with humans too?). Clearly, too, GPT was helped by the facts that human mathematicians had wasted most of their time trying to prove Erdös right rather than looking for a counterexample, and that, even if they did look for a counterexample, they’d need to be experts in algebraic number theory to find this one, which hardly any discrete geometers are. So, maybe that suggests that AI, right now, is “merely” picking various medium-hanging fruits that human mathematicians missed for contingent reasons? With emphasis on the “right now.”

In a companion paper, OpenAI helpfully included commentary from Timothy Gowers, Noga Alon, Will Sawin, Daniel Litt, and many other experts, reflecting on the breakthrough, the path that GPT took to get to it (which can actually be seen by examining its chain-of-thought), and what this might mean for the future of mathematical research.

I heard the news maybe an hour after it broke, when some UT grad students came to my office to tell me. For what it’s worth: these students were morose, musing about how everything might soon be over for young scientists and mathematicians like themselves. I don’t know whether they’re right, but I feel like I should tell the truth about what their reaction was.

[Update: News has been coming faster than I can write about it, but today we learned that another important conjecture of Erdös has been refuted. Erdös and Szemerédi’s strong sumset conjecture over R, from the 1970s, had said that, if A is a finite set of real numbers, then either |A+A| or |A×A| must be at least |A|2-o(1). In this case humans, including the aforementioned Sawin, did almost all of the work of constructing the counterexample, but they were directly inspired by GPT’s earlier refutation of the unit distance conjecture. It remains open whether such a counterexample exists where A is a set of integers.]

Then, a few days later, a team at DeepMind, including my UT Austin colleague Swarat Chaudhuri, announced that they were able to use a system called AlphaProof Nexus to settle nine more (!) Erdös problems, many of them in additive combinatorics, along with miscellaneous other open math problems. Notably, in this case the AI also fully formalized its proofs in Lean.

And then, just today, Jelani Nelson alerted me to a new CS theory paper, which solves a longstanding open problem about electrical flows on graphs using a proof from GPT5.5Pro.

It seems to me that we’re now over the top of this particular rollercoaster, and it will keep accelerating until we reach the bottom, wherever that might be. I don’t know whether to hope or dread that solutions to P versus NP and all our other great problems will be included in the ride—that our role, as human mathematicians, will be reduced to (at most) deciding which questions we find interesting and then understanding AI models’ answers to those questions.

But maybe that won’t happen. Maybe the new AI mathematicians will soon hit a wall, because they lack the uncomputable quantum gravity microtubules of Penrose and Hameroff, or some other magic human ingredient. The fantastical thing is that, one way or the other, we’re going to find out empirically before very long.


Readers may have also seen the news that multiple prizewinning entries in a short fiction contest called the Commonwealth Prize, give overwhelming indications of having been written by AIs. As Kelsey Piper puts it:

There are, let’s say, also some noticeable similarities in the prose style between the winning stories that were flagged for AI use. AI chatbots love metaphors and similes, and they often spit out ones that sound vaguely pleasing but are logically incoherent or ascribe properties to things that don’t make sense.

“The Serpent in the Grove” gave us, “The girl smiled like sunrise over a sink.” “The Bastion’s Shadow” says, “She carried it now in her bag, heavy as a charm.” “Mehendi Nights” describes something as “swaying against plaster like a warning bell.”

The Commonwealth Foundation, whose judges chose these stories, hasn’t exactly covered itself in glory—saying, on the one hand, that it strictly forbids AI use but on the other, that it will continue taking authors at their word that they didn’t use AI, no matter the immensity of evidence to the contrary. As many others have pointed out, judges more versed in AI would’ve ironically been better placed to notice the signs of its use.

If only there were some sort of automated way to detect AI-generated text. Someone should really get on that problem, don’t you think?


But maybe we should just throw in the towel—as some of my colleagues have already done in the context of undergraduate projects? Maybe we should simply say that a good story is a good story, regardless of what manner of entity produced it?

As it happens, just last week I read my very first AI-written story that affected me as a story, to the extent that I wanted to read it more than once. This happened when I gave GPT5.5Pro the following simple prompt:

Write me a story about the most ancient Israelites that’s riveting like the stories of the Bible but that’s also consistent with all of the archeological evidence.

You can read the result here. One of my Facebook friends called it “disturbingly good,” and whatever the problems with the piece, I share that feeling. Of course, I’m well aware that GPT could easily generate a thousand stories like it—sampled from the same probability distribution—and then I could even do statistics on which tropes were the most common. This makes it feel silly to overindex on the first story that happened to be output, and yet somehow I did.


I feel like at this point, both the prophets of AI utopia like Ray Kurzweil, and of AI doom like Eliezer Yudkowsky, could be forgiven for asking: dude, will you listen to us YET? Do you still find it prudent to call this new form of terrestrial intelligence a stochastic parrot, a laughable fraud, or a fad that’s about to go away? Fear it all you want, hate it even, but at least respect it!

Which brings me to the other big AI news from the past week, namely that Pope Leo released his first encyclical, which is entitled “Safeguarding the Human Person in the Time of Artificial Intelligence.” I read it and … well, I certainly agreed with the theme that such a world-changing technology needs to be developed for the common good (as the Pope would have it, like the walls of Jerusalem), rather than for the profit or vanity of any one individual or company (in his analogy, like the Tower of Babel). I had quibbles with some of the other parts. Zvi Mowshowitz, as he often does, had a superb paragraph-by-paragraph analysis. Amusingly, there are indications that parts of the encyclical were written by AI.

To me, though, maybe the most notable part was that Chris Olah, who leads Anthropic’s interpretability team, was standing next to the Pope at the ceremony, and delivered his own remarks. I felt like Chris, who I met even before Anthropic existed, was a non-obvious yet inspired choice here, one of the rare figures in frontier AI whose technical and moral authority are both completely unimpeachable by anyone.

And so, at this momentous era for the human project, and on no less of an authority than that of the Vicar of Christ himself, the Supreme Pontiff and the Successor of Peter, I hereby throw myself on the wisdom and mercy of … uhh, I guess, Chris Olah and his team at Anthropic.

Chris, if I am soon to share the earth with entities that can prove the Riemann Hypothesis and solve quantum gravity after 30 seconds of thought, then may you understand those entities well enough to cause them to be nice.


Endnote: I should have foreseen, but didn’t, that the comments on this post would be dominated by people looking for ways to minimize whichever specific AI accomplishments I blogged about. Thus, it turns out, the ability of AI to solve Erdös problems just demonstrates that Erdös’s problems were never “serious” math in the first place—nothing like algebraic geometry or Grothendieck-style theory-building, which remain untouched. Likewise, the story I shared was obvious AI slop.

I had taken it as obvious that, when assessing AI’s impact on the world, one needs to look at least somewhat into the future: to remember where things were four years ago, compared to where they are today, and at least try to draw a straight line through the data, if not the exponential that seems to fit better.

Does anyone seriously doubt at this point that major open problems in algebraic geometry and other “Grothendieck-friendly” areas of math will fall to future AI models? Or that AI-written stories will improve, not only to win literary awards from AI-naïve judges, but to avoid the features that commenters here are complaining about? And that, whenever that happens, there will be new confident reasons not to care immediately offered up in comment sections like mine?

Apparently people do still doubt—hence the throwaway remark in my post about Penrose and Hameroff and microtubules. If not that or something like it, what exactly do they think the ceiling will be, and why?

Recently (I should have mentioned this before), I came across what I consider one of the greatest social experiments of all time, one that illuminates people’s reactions to every AI advance. A Twitter/X user named SHL0MS displayed the following AI-generated fake “Monet painting,” and asked people to explain what made it worse than real Monet paintings:

If you haven’t seen this yet, I recommend that you try the exercise yourself before reading further.

As it was, numerous art aficionados responded at length, savaging the flat, lifeless, uncreative AI slop, the emotionless composition, the missing spark, the lack of tranquility, the harshness, the lack of depth and symbiosis, and on and on and on.

Only after they had all said their piece did SHL0MS reveal that this is an actual Monet painting.

222 Responses to “Dispatches from the possibly last days of human relevance”

  1. Prasanna Says:

    Scott,
    ” I read it and … well, I certainly agreed with the theme that such a world-changing technology needs to be developed for the common good (as the Pope would have it, like the walls of Jerusalem), rather than for the profit or vanity of any one individual or company”

    Well, add nation-state to this line. And good luck with that. We humans are a product of evolution and the baggage of limitations that accompany it, so the only way this outcome can happen for a powerful technology is for the technology itself to take over humans. As ironical as it sounds, AI neeeds to literally become God for this outcome to stand a chance

  2. Daniel Says:

    I think I’d characterize the voice of your story as distinctively “tumblr”. (In particular it is very similar in style and content to things by Eurydice Lives, with her pacing/cadence vices exaggerated.) I guess there are probably a lot more tumblr posts than literary magazines in the training set.

  3. Hyman Rosen Says:

    The vast majority of humans cannot prove the advanced mathematical theorems that mathematicians routinely generate. For the most part, they cannot even understand what those theorems mean. And yet, somehow, those humans don’t feel that they lack “human relevance”. For that matter, they cannot do arithmetic as quickly as a computer, run as fast as an automobile, fly through the air like a plane, break rocks like a jackhammer, or do many other things as efficiently as the devices that a tiny number of humans have invented and a slightly larger number of humans know how to build. And yet they still muddle along, mostly content with their lives. Human relevance is not tied to the most ambitious thing that the best human can do. There isn’t even such a thing as human relevance. Humans don’t need to be relevant to anyone but themselves and to their family and neighbors. I’m retired. I sit around and play video games, read books, go to movies, and do things my wife tells me to do. I have been moderately useful in my career to those who have employed me, I have helped raise a son who is so far successfully launched in life, and I feel fine with that, despite not having invented anything, not written any novels, nor set any athletic records.

  4. Not Owl Sowa Says:

    “the possibly last days of human relevance”

    FFS, can we stop with this kind of hype? It is a great achievement of AI. No doubt about that. But the anxiety from Gowers and the related “problem-solving culture” crowd is honestly overblown.

    Note that the AI solution to the planar unit distance problem relied heavily on deep theories developed by algebraic number theorists, going back to Gauss, Artin, and many others. Mathematics is not just a collection of isolated puzzles proposed by some Hungarian guy. It is centuries of accumulated theory building.

    Hungarian-style combinatorialists ignored large parts of mathematics for decades, and now they are paying the price. The rest of mathematics can use AI as an amplifier rather than a replacement, and will remain highly relevant.

  5. Sniffnoy Says:

    You know, I think if I were just starting out, I might not mind it so much — LLMs could be part of my method, right?

    The problem is, I’ve got a huge backlog to write up (especially due to my semi-outsider status, where I have to get this done in my free time), and what I worry about is that someone using AI will, like, go scooping that. :-/ You don’t have to worry about this so much with people due to the community nature of mathematics — people have an idea what other people are doing and often keep each other informed by less formal means before they go writing up papers. Yes, sometimes it happens due to people working outside their normal area or people just not being very communicative, but to a large extent people working on similar things keep in touch and may even talk to each other about how credit should be divided up. An LLM knows none of that! (I guess, if it’s read my Dreamwidth blog and website, it may know that I’ve informally posted there results in those places that I haven’t yet written up as papers; but current LLMs often don’t cite their sources in the final result unless you check the chain of thought. Something that needs fixing.)

    …you know what I maybe should try to use an LLM for, though? Making sense of Achim Flammenkamp’s thesis on addition chains. I have an English translation provided to me by Neill Clift; it being in German isn’t the problem. The problem is that it just isn’t written in the usual definition/theorem/proof style of mathematics, and large amount of what he’s done is just left implicit. It’s so utterly opaque, I’ve bounced off it a number of times, maybe if I sat down with it for a few weeks I could figure it out but I’ve never had the patience to do that. Maybe I ought to see if Claude can make sense of it. 😛 (Yeah I know this paragraph is not going to make any sense to anyone who hasn’t studied small steps in addition chains and thereby encountered Flammenkamp’s thesis…)

  6. Bud Says:

    I find this existentially upsetting. On the one hand we’re going to see results and advancements come thick and fast and to god knows what end. On the other hand, I feel like everything I’d ever worked to understand and to is now moot.

    The basic question is: what do we do with ourselves when our intelligence is literally unnecessary? I see myself shriveling into nothing.

  7. Amazing: Erdős’ Unit Distance Problem was Disproved! It was achieved by AI! | Combinatorics and more Says:

    […] Github. The result is reported and discussed on the Geomblog, computational complexity, and Shtetl Optimized; There is an interesting commentary by Erik Hoel. Anthropic reported being able to autonomously […]

  8. Vamsi Says:

    The performance of AI appears to be uneven across areas. In analysis (I work in geometric analysis – the application of PDE to differential and algebraic geometry), it is super human at examples and counterexamples, and possibly at proving concrete lemmas (or perhaps at even modifying a few of them in a given strategy/paper) using clever tricks (if IMO winning school students knew PDE, you may expect this sort of intelligence from them). However it is not good at coming up with global strategies. So it is locally great but not globally (but it appears to globally great at some algebra/discrete things going by the grapevine). My theory is that world models need to be integrated with LLMs to become globally good at analysis (also maybe formalisation in lean to act as a verifier). Such a hypothetical AI would probably replace most of us. Sure it may not replace Terry Tao but the future Terry Tao will be scared of job prospects in maths anyway.

  9. Danylo Yakymenko Says:

    then may you understand those entities well enough to cause them to be nice

    I’m not sure about those AI entities, but a lot of the current entities that have the power to control AI development and usage don’t show signs or intentions of being nice (except to their temporary peers). I believe that if AI were to figure out how to prolong human life indefinitely by sucking the blood of newborns, then many of the entitled ones would happily take this sacrifice.

    Although, I think it is good to see how the humankind develops through this groundbreaking technology. It is like a litmus test for many. But it may be a false start for those who think they will end up at the very top of the new societal order. The real God may not have a transformers heart.

  10. Avi Says:

    Re the ancient Israelite story, I have to completely disagree with your assessment.

    First, it is utter AI slop in style, impossible to read to me and that is independent of any particular theme or detail in the story. I hear that some AI models in spring 2026 can produce fiction that is actually worth reading, but this model clearly can’t.

    Second, already in the first line there is a pretty obvious historical error (“before a house of cedar stood for any god”), people in this time and place were not atheists, they did have temples, the AI is taking a lack of archaeological evidence for which deity was worshipped and misunderstanding it as the people at the time not having decided which deity to worship.

    Third, what appears to be the main plot point (re Merneptah’s stele) seems to be handled in a very heavy-handed and awkward way.

  11. Lucius Says:

    You don’t have to passively throw yourself at the mercy of Chris Olah and leave it all up to him and his team you know. You can join us in trying to figure out how on earth the AIs work.

    Maybe you’re doing that already, but I point it out just in case you’re not.

  12. Tatum Says:

    I think @Prasanna misses something important with their comment:

    >As ironical as it sounds, AI neeeds to literally become God for this outcome to stand a chance

    The game theory rather compels the de facto atheist countries of the world to act on their natural enmity towards God by going full King Herod on any nascent superintelligence project. Given the inscrutability and illegibility of any attempted RSI process, that in practice means threatening to airstrike the servers of any company that is aiming to create an automated AI researcher. Current timelines suggest that this threat must be made and possibly acted upon by 2028 at the latest.

    (To set the tone for any military operation that unfortunately results from such a position, the appropriate background music might be Zankoku na Tenshi no Tēze.)

  13. Ben Green Says:

    Scott, I’m pretty sure this wasn’t solved by GPT 5.5Pro. It was an internal model at OpenAI.

  14. Edo Blaauw Says:

    Well you Will become anagolues to body builders. Showing your mental skill to asking difficult questions with AI’s being the judge saying if you are correct or not and people without the faintest idear why you are correct cheering you on because of your smartness. People Will play good money for that. It Will be kind of a carnaval act but that’s that.

  15. nesuno Says:

    really? maybe AI beats Erdös, Euler and Ramanujan (solves problems), but does that mean it can replace the likes of Dedekind, Gödel, Grothendiek or Hilbert (theory building)?

  16. Jon Says:

    An unexpected outcome of all of this seems to be that the increase in respectability/prestige of combinatorics over the past decades might see a sharp reverse! The culture of problem lists and opennness to all questions has left an orchard to be plundered (this is not a criticism).

  17. Jean Abou Samra Says:

    Is it really GPT5.5Pro? The announcements mention “an internal AI model”.

  18. Chris Says:

    I think you are doing a little bit of cherrypicking here. At the same time these impressive math results were being generated by AI, reports also came out that AI use for coding has not provided the productivity gains predicted.

    What explains this?

    To me it seems that math proof is actually a very specialized task with well defined criteria for success. This is something we might expect a computer to be good at. Indeed, automatic theorem provers have been around for a long time, and nobody ever said these had superintelligence.

    Unfortunately, most tasks, including programming, lack clearly defined criteria for success. In fact, for programming, I would suggest the main problem is actually figuring out what you want it to do in the first place.

    That being said, I would love to try these models for proving some conjectures I am working on myself. As far as I am aware, these models are not available for public use. Am I mistaken?

  19. Jeremy Says:

    For the moment, successful AI arguments remain short and concentrated in areas with a lot of literature. I am curious about *when* AI will be able to construct the coherent, long-form papers that are most relevant to theory-laden fields. I don’t doubt that it will, but I wonder whether the obstruction to doing so is the same obstruction to the machines being able to autonomously generate income. Is crafting a creative popular video game or novel, that hangs together over its whole context without human input, going to prove harder than generating a long Annals paper in homotopy theory?

  20. Liron Says:

    I wonder if it might soon be the case that in practice “P = NP + AI”, in the sense that AI will be able to find a solution, approximation, shortcut, or satisfactory workaround to any NP problem we care about. Then our only hope is that taking over the world is closer to PSPACE-hard and AI algorithms won’t reduce problems so easily there.

  21. LK2 Says:

    Hi Scott,

    impressive performance of GPT, but you seem to mention mostly successes in discrete/combinatorial math. Have LLMs scored similar successes also in “continuous” mathematics like, I dunno, algebraic/differential geometry, analytic number theory, analysis, and so on?
    If not, could be that discrete math better adapts at their reasoning model?

    Last comment/question: great that they solve these things, but since the LLMs do not have “intentionality” but -for now- only humans spontaneously (whatever it means) interrogate themselves on certain problems, it seems to me that humans are still essential for 1) asking 2) checking. After all, a non-mathematician would hardly come up with an Erdös problem…
    What do you think? Is my line of thought faulty for some reason?

    Best,
    K
    (and congratulations for your recent prizes and recognitions: well deserved).

  22. wb Says:

    It turns out that math is easier than driving a car …

  23. Scott Says:

    Ben Green #13 and Jean Abou Samra #17: Oof, sorry about that! Fixed.

  24. egan Says:

    >>>It seems to me that we’re now over the top of this particular rollercoaster, and it will keep accelerating until we reach the bottom, wherever that might be.

    For the moment there is no AI breakthrough in theory-heavy math fields (arithmetic geometry, algebraic geometry, etc). Discovering a solution to a problem in combinatorics is not the same thing than inventing Perfectoid spaces or Condensed mathematics.

    >>>AI chatbots love metaphors and similes, and they often spit out ones that sound vaguely pleasing but are logically incoherent or ascribe properties to things that don’t make sense.

    Famous opening line of Neuromancer : “The sky above the port was the color of television, tuned to a dead channel”.

  25. Mike-e Says:

    Hyman #3

    “ The vast majority of humans cannot prove the advanced mathematical theorems that mathematicians routinely generate. For the most part, they cannot even understand what those theorems mean. And yet, somehow, those humans don’t feel that they lack “human relevance”

    Things work this way: a lot of humans need to be born in order to create more specialists in a given field. It’s a statistical game. If human population was capped a 2 millions instead of 8 billions, chances of improving the 100m dash or coming up with general relativity falls sharply.
    There’s also the fact that since each human life is finite, one can’t possibly focus on everything- the vast majority can’t read academic papers, but similarly the vast majority of academics can’t do plumbing and cut hair.

  26. exponent Says:

    Sawin’s improvement to the exponent has been subsequently improved a few times with the assistance of ChatGPT 5.5 Pro (the public version).

    https://mathoverflow.net/questions/511514/what-is-the-unit-distance-exponent

    Sawin did not use AI in his paper intentionally. But he anticipated that the exponent would be improved and suggested a number of possible ways.

  27. Mike-e Says:

    Jeremy
    “ For the moment, successful AI arguments remain short and concentrated in areas with a lot of literature.”

    That’s the thing with math, the space of open problems is nearly infinite, and AI will solve first whatever problems humans have been focusing on, historically, by definition, since LLMs are based on training from existing knowledge.
    It’s possible AIs will eventually start taking to explore sideways in new directions, using their own results as a basis to move further and further away from human’s original areas of interest. At this point human researchers won’t be able to keep up because new knowledge will increase faster than a team of coordinated humans can keep up with.

  28. Mike-e Says:

    Human activity will soon be entirely focused on:
    1) coming up with ways to build data centers more efficiently
    2) coming up with ways to destroy data centers more efficiently

  29. Joshua Zelinsky Says:

    Avi #10, Second, already in the first line there is a pretty obvious historical error (“before a house of cedar stood for any god”), people in this time and place were not atheists, they did have temples, the AI is taking a lack of archaeological evidence for which deity was worshipped and misunderstanding it as the people at the time not having decided which deity to worship.

    I interpreted that line as not referring to people being atheists but the idea that more complicated temples which required grand construction didn’t yet exist. That seems consistent with some of the later parts for example where the story reads:

    > At the standing stone, the captain ordered the offerings kicked aside.

    > “This god of hills has no house,” he said. “How shall he protect a people who have no walls?”

    > Oren answered, “Perhaps a god without a house is harder to burn.”

    And later:

    > She looked over the hills. There was no palace. No army. No cedar temple. No scribe to make their grief magnificent.

    For it is worth, I was when I first going to read this story say it showed that Scott was easily persuaded and that it range as sterile and as hollow as so many other AI stories before that. But I found it touching. However, I’m also someone who is really easily emotionally moved but what my spouse called “rousing songs of humanism” or the like. I’ve rewatched the first episode of Strange New Worlds and cry every time at Pike’s speech. So I’m really not a good judge. But this is at least the first time I’ve seen an AI do it. I suspect my ethnic heritage is also playing a role here. But this is also just genuinely much better done than AI stories even six months or a year ago.

  30. Karim Says:

    Though it was an internal model at OpenAI, some people on X seemed to have gotten 5.5 Pro to solve it with some minor additional prompting too.

  31. Scott Says:

    Lucius #11:

      You don’t have to passively throw yourself at the mercy of Chris Olah and leave it all up to him and his team you know. You can join us in trying to figure out how on earth the AIs work.

      Maybe you’re doing that already, but I point it out just in case you’re not.

    For the past year, I’ve been running the Theory and AI Alignment Group at UT Austin, generously supported by Coefficient Giving. We’re trying to prove theorems relevant to out-of-distribution generalization, watermarking and backdoors, and (yes) interpretability. Here I’ve been inspired by the stuff I and others did in OpenAI’s now-defunct Superalignment team in 2022-2024.

    On the other hand, our group at UT almost certainly has no comparative advantage at looking inside models empirically, the way Olah’s group at Anthropic famously does. We’re complexity theorists, ready to help if we can.

  32. John K Clark Says:

    I think it very unlikely that Penrose and Hameroff are correct, it seems to me that the interior of neurons are too hot and chaotic for quantum entanglement, but if I’m wrong and they’re correct then that would mean it would be much easier to make a large fault tolerant quantum computer then had previously been thought; and we’d still be living through the last days of Human relevance.

    John K Clark

  33. Scott Says:

    Liron #20:

      I wonder if it might soon be the case that in practice “P = NP + AI”, in the sense that AI will be able to find a solution, approximation, shortcut, or satisfactory workaround to any NP problem we care about.

    I doubt that this is true, if only because NP problems include all the cryptographic problems that were specifically constructed to be hard (e.g. bitcoin mining, LWE). And also, while I can imagine an AI that proves P≠NP and the Riemann Hypothesis, it’s much harder for me to imagine an AI for which no natural open math problem (even, let’s say, among those that ultimately do have reasonable-length proofs) ends up being too hard.

    On the other hand, yes, something like what you say will almost certainly become increasingly true as AI advances, at least outside the realm of cryptography. Indeed, that’s a huge part of the point of AI!

  34. Scott Says:

    wb #22:

      It turns out that math is easier than driving a car …

    Uhhh, AI has been driving cars for years now! I ride Waymos pretty regularly here in Austin. The main obstacles are now in

    (1) manufacturing enough cars cheaply enough (especially when they need onboard LIDAR), and especially

    (2) getting past regulatory blankfaces who are congenitally unable to think statistically (“we’re already getting into 90% fewer crashes than human drivers!” “AHA, then you admit you’re not 100% safe!”).

    Certainly for driving in well-mapped urban areas (i.e., the majority of all driving), the main issue is no longer about the underlying technology.

  35. Nilima Nigam Says:

    Indeed, I’m curious about where we’ll end up on this roller-coaster. For now it is clear we’re on some gravity-defying, sometimes scary predetermined track (in your metaphor). I don’t think anyone’s really ‘steering’ this collection of loosely-jointed wagons, but clearly all are being propelled somewhere.

    In other news, I read the phrase ‘meat computer’ for the first time last week.

  36. wb Says:

    Scott @34

    >> AI has been driving cars for years now

    I am aware of that, but Waymo needs Lidar and very well mapped streets,
    Tesla needs several cameras and still cannot really get there yet.

    Humans can drive in cities they have never been to before, and also off-road etc.
    In short, AI is already beating professional mathematicians (and go and chess players) but not yet human drivers who navigate roads with two sub-optimal cameras …

  37. Tomas B. Says:

    Everyone is having a crisis of meaning. I sympathize. At risk of advertising, I wrote this when my own obsolesce hit me emotionally: That Mad Olympiad

    Ilya Sutskever once said something like, “if you get your self worth purely from your intelligence, you’re going to have a bad time.”

    There are far bigger issues facing us. But there is something lost here that will never return.

  38. Philip Weiss Says:

    @Prasanna #1

    Maybe what we need is… Model UN

  39. Avi Says:

    Joshua #29,

    Canaanites definitely had temples in this period, and a couple minutes of googling reveals them to be large and well-constructed:
    https://commons.wikimedia.org/wiki/Category:Tel_Hatzor_-_Canaanite_temple#/media/File:Tel-Hazor-1859.jpg
    https://www.afhu.org/2020/02/18/hebrew-university-team-unearths-canaanite-temple-at-lachish/

    I definitely see the potential attractiveness of a well researched story in a biblical setting – in a way, it’s like a new book of the bible has been suddenly discovered, giving a new window onto a culture and historical moment that have a notable place in many of our minds. But I just couldn’t get past the AI-ness of this story.

  40. Dániel Says:

    > Clearly, too, GPT was helped by the facts that human mathematicians had wasted most of their time trying to prove Erdös right rather than looking for a counterexample, and that, even if they did look for a counterexample, they’d need to be experts in algebraic number theory to find this one, which hardly any discrete geometers are.

    I am one of those discrete geometers who wasted quite some time on this problem, and I’d quibble with one minor point. Looking for a proof seemed so hopeless, and finding a counterexample so much more realistic, that I devoted all my effort to finding a counterexample. I suspect I wasn’t the only one. The second part of the sentence describes me perfectly: I did not have a chance because I wasn’t up to par in algebra.

  41. Gabriel Gaster Says:

    And now the sum product problem must be added to this rapidly growing list, since it directly uses a line of attack from the LLM paper.

  42. Ajit R. Jadhav Says:

    Dear Scott,

    I read your above post with interest. Looks like it was a call to (as many)^N souls as possible for whom as many as N Bells toll!

    Too poor a reach if most of the souls to whom you could even manage to reach out even in theory were mostly mathematicians (say with Erdos number much less than the long ago revered number 6), or similars from Industry and all.

    As to the more immediate results of interest about which you post:

    What is the exact cost in terms of “weight” of gold (or if you prefer even now, the US $), for all the queries (“prompts”) by you all? I mean, including those by the “agents” and all, set up by you all?

    What would be the cost if there were to be no subsidies at all? Say those offered by the promoters of these machines?

    [Recall: Projections of revenues of those who offer these services; the revenues accruing to them as of now; the special discounts to ladies and gentlemen of some National thing of which you only recently were elected to be a member; and similar considerations.]

    Regardless of the above square-bracketed comment, I remain concerned about the absolute cost to the end user, e.g., you.

    Should you have a reply to offer covering all the important points, please do post it here. I should be checking this blog either the day after tomorrow, or soon thereafter.

    Best,
    –Ajit
    [Written rather hurriedly and on the spur of the moment, more often “as usual” than not.]

  43. David Speyer Says:

    The new sum-product result (I assume you mean https://arxiv.org/pdf/2605.28781) is a human generated result, although closely inspired by the unit distance proof. It is certainly part of the same story, but the moral here seems to be “very smart humans who read AI output and learn from it can still make valuable contributions”. To whatever extent you think we are in the last days of human relevance (for solving math problems), this paper should make you think those days might last a little longer.

  44. Simon Lermen Says:

    Comment #6

    > The basic question is: what do we do with ourselves when our intelligence is literally unnecessary?

    Hey Bud,

    You know there is a huge open field known as AI alignment that desperately needs smart mathematicians working on stuff like interp. This is a field that is pretty hard to automate, why don’t you tell some of your maths friends to take a swing at it, perhaps starting with some of the stuff Neel Nanda has been working on? You can actually make a difference there and we can’t just hand it off to AI before they become much more dangerous.

  45. Harvey Lederman Says:

    Bud #6, see earlier expressions of depression (but also optimism?) on this site… https://scottaaronson.blog/?p=9030

  46. GeorgeM Says:

    Chris #18

    > That being said, I would love to try these models for proving some conjectures I am working on myself. As far as I am aware, these models are not available for public use. Am I mistaken?

    The specific unit distance problem was solved with an internal model at OpenAI, which I guess may or may not eventually be released. Surely something of similar power will be released at some point.

    But Erdos #1196 was solved* using GPT 5.4 Pro, which is available publicly, at $200 / month.

    * see this comment https://www.erdosproblems.com/forum/thread/1196#post-5361

    > To me it seems that math proof is actually a very specialized task with well defined criteria for success. This is something we might expect a computer to be good at. Indeed, automatic theorem provers have been around for a long time, and nobody ever said these had superintelligence.

    It’s using a computer in a very uncomputery way though! Yes, it solved a problem that was very objectively verifiable, and it seems the verifiability is important to LLM’s developing mathematical capabilities.

    But the way computers were taught this skill was by teaching them how to read and write the nuances of natural language, then how to use natural language to support reasoning for themselves. So we’ve gone from objective computers to fuzzy language, then back to objective mathematics. That path kills – for me at least – any feeling of the kind “of course computers can prove theorems”.

    I’m not aware of any automatic theorem provers that have got close to being able to solve this level of problem from scratch. But I’d be interested to be corrected on that!

  47. Nobody Important Says:

    Regarding the “last days of human relevance”, I think we are a long way away from that. When was the last time you saw GPT 5.5 pro fix a flooded sewer line while being inundated by water, mud, and sewage? Continuing to function while covered in muck and shit is a lot harder, apparently, then solving Erdös problems.

  48. Joshua Zelinsky Says:

    Avi, #39,

    Canaanites definitely had temples in this period, and a couple minutes of googling reveals them to be large and well-constructed:

    Sure. And that’s a valid complaint. That’s very different than your statement that “Second, already in the first line there is a pretty obvious historical error (“before a house of cedar stood for any god”), people in this time and place were not atheists, they did have temples, the AI is taking a lack of archaeological evidence for which deity was worshipped and misunderstanding it as the people at the time not having decided which deity to worship.”

    So, your own initial reaction was about as bad as the AIs here.

    I definitely see the potential attractiveness of a well researched story in a biblical setting – in a way, it’s like a new book of the bible has been suddenly discovered, giving a new window onto a culture and historical moment that have a notable place in many of our minds. But I just couldn’t get past the AI-ness of this story.

    Sure. I have some trouble with that, and obviously a lot of people have that reaction not just for stories but for other things which feel like they should have an emotional connection to humans. My own reaction to the story as I said may be in part due to my own idiosyncrasies about the topic. And now that I think about it, I’ve actually written a short story which in some respects is similar in topic but with a jump to the future (https://joshuazelinsky.dreamwidth.org/5749.html ). There’s now a part of me that is tempted to go ask AI to write stories with the same premises as my own stories and see if it does well, and most of me is violently rebelling at the idea in deep disgust.

  49. Martin Aulbach Says:

    Funny that you should mention Penrose’s quantum microtubules, Scott. Recent research by renowned neuroscientists (Oxford-based Prof. Morten Kringelbach) indicates that there are quantum-like processes going on in our brains:

    https://www.biorxiv.org/content/10.1101/2025.10.02.680057v1
    https://academic.oup.com/book/62453/chapter/556522238

    If true, that would be encouraging news, as AI is definitely *not* going to catch up on this anytime soon. What do you make of this, particularly with your background in computational complexity theory? Could such quantum effects possibly explain the Dual Process Theory of psychology with a faster, unconscious process and a slower, conscious process? I wonder if these two processes could be somehow understood as quantum and classical computation, respectively, with free will possibly hiding out somewhere around there too.

  50. Balazs Says:

    > getting past regulatory blankfaces who are congenitally unable to think statistically

    That might be true, and even if I am able to think statistically (or at least I like to think so) here I am firmly on the same side as the blankfaces i.e. default to policy of AV restriction. I would argue, statistical thinking here in the narrow sense has a quite a few blind spots. Handing over public spaces to autonomous vehicles / robots (and by extension to the companies who develop them) isn’t only a question of accident statistics. There are emergent risks of governance eg. regulatory and infrastructure capture (essentially, AV companies reshaping laws and infrastructure to benefit their business model).

    You might say that AV companies business model aligns with human interests, but I would argue this is only temporary and by coincidence. Eventually, AV companies profit by maximizing rides. Humanity profits from efficient travel (minimizing unnecessary travel, using efficient modes of transport).

  51. Raoul Ohio Says:

    Scott #34, Balazs #50

    blank faces are denying us the wonderland of driverless cars? it woud be interesting to see how if voters agree.

    90% fewer crashes? is this number from peer reviewed research, or from marketing hype?

  52. Vladimir Says:

    > […] our role, as human mathematicians, will be reduced to (at most) deciding which questions we find interesting and then understanding AI models’ answers to those questions

    One of my favorite Asimov stories (“Jokester”) expands on that premise.

  53. Scott Says:

    Raoul Ohio #51: Yes. Blankfaces, who you apparently support, are denying us driverless cars—for example, by preventing the expansion of Waymo to Boston and NYC and DC, even though it’s been great in SF and Austin and everywhere else it’s available. This is why many of the key decision makers lack direct experience with a technology that would save 40,000+ lives every year if they allowed it to.

    80-90% fewer crashes is now a robust result from more than 200 million miles driven by Waymos. The rare times there are accidents, it’s almost always the human driver who’s at fault.

    Of course, for those like you, no amount of data will ever suffice, just like it never did for those who killed nuclear power and thereby set the stage for today’s climate catastrophe. The 200 million miles will become 2 billion miles and you’ll still say it’s all marketing hype and more data is needed. That being so, why pretend this was ever about data in the first place?

    One more thing, though: I have no idea why you imagine I’d treat whether voters agree (!) as relevant to the truth or falsehood of what I just said. A large fraction of voters also want mRNA vaccines to be banned. They’d rather their own families die than have to break their own epistemic closure, to think quantitatively rather than via viral anecdotes on TikTok. Have I ever had the slightest problem admitting that voters can be complete fucking idiots?

  54. Scott Says:

    Vladimir #52: My daughter and I are now reading Asimov’s 1953 The Caves of Steel together; I’d last read it when I was an adolescent myself.

    When you get past all the anachronistic stuff — the 1950s obsessions and cultural assumptions projected forward a thousand years that were already obsolete by the end of Asimov’s life, let alone the 2020s — there’s also lots of stuff that Asimov correctly foresaw, maybe that was even easy to foresee. The whole background to the novel’s murder mystery is a human population that’s angry and fearful about the wider deployment of robots, sometimes inventing other reasons but really worried about being outcompeted at their jobs.

  55. Scott Says:

    Balazs #50: So then, would you accept full self-driving if it were a government project, rather than something by private companies begging the government to let them operate?

    Which previous technologies would you have similarly banned with the same precautionary logic: airplanes? roads? human driving?

  56. Anon Says:

    It is a very exciting and high anxiety time to be alive.

    AI will be transformative, no doubt about it. But so far these systems do not show ability at the highest human level and many experts (including Ilya, Demis, Yaan, …) believe there are some fundamental issues with these systems, e.g. in an interview Ilya mentioned that these models repeated keep making simple mistakes repeatedly (fixing bug with a new bug, then fixing that bug and causing the original bug, and repeating the pattern) and scaling doesn’t seem to be improving this. Even Ilya who used to be a hyper optimistic that these systems will just scale to be as good as humans have reversed his opinion and considers them to the current systems to have some fundamental short comings.

    The truth lies somewhere between the AI hypers and the AI skeptics.

  57. Hyman Rosen Says:

    Speaking of human relevance, have you commented yet on the recent thing about the quantum supremacy claim that was overturned by using tensor networks on ordinary computers?

  58. Dacyn Says:

    Balazs #50:

    You might say that AV companies business model aligns with human interests, but I would argue this is only temporary and by coincidence. Eventually, AV companies profit by maximizing rides. Humanity profits from efficient travel (minimizing unnecessary travel, using efficient modes of transport).

    AV companies’ business models align with their customers’ interest just as much as any other company’s business model does. The company wants to maximize revenue while minimizing cost to provide the product, while the customer wants to maximize value received from the product while minimizing the cost paid to the company (which is equal to the company’s revenue). For the most part, sales happen when transactions are beneficial to both parties, which can happen if and only if the value the customer gets from the product is greater than the amount it cost the company to provide the product.

    To put it simply, the customer’s protection is that he can choose not to buy the product if it isn’t a good deal for him. He doesn’t need the company’s interests to “align” with his own interests, and in fact as described above their interests are exactly opposed in one important respect (the amount of money paid from the consumer to the company). I don’t see how any of this is any different when we’re talking about autonomous vehicles versus anything else, now or in the future.

  59. Scott Says:

    Hyman Rosen #57: Which one? That sort of thing has happened often over the past 7 years, just like proposed cryptosystems often get broken. But the best current random circuit sampling experiments, especially the recent ones from Google and Quantinuum, seem well out of reach of even the most optimized tensor network simulations that are known.

  60. Chris Says:

    @GeorgeM

    Thanks for pointing this out! I take a little bit of issue with the contention that the computer did this in a very uncomputery way. In the google case at least, the computer seems to have used at least something close to the architecture of the other Deepmind approaches. Please correct me if this is wrong or oversimplified: they used a reasoning neural network iteratively with a formal proof verifier. The proof verifier was a way to prune the search space so that it became more tractable. They used a model trained on existing proofs to give it the various steps to try. This seems rather computery.

    I want to emphasise, however, that I actually really like this approach! I want to be able to use it for my own research. I also agree with Scott that we can infer that a lot more proofs will come from this. I myself am looking forward to a result in computational complexity theory.

    On the other hand, I have a concern about the ChatGPT 5.4 Pro approach. The issue is that I currently don’t trust it. Hence, it may very well produce a correct proof, but without confidence of it being true, will I have the willingness to expend potentially a large amount of my work day trying to verify that it is true?

    @Scott: I admire your enthusiasm about AI. But I have to quibble with something. It has to do with your art example.

    The reason art is beautiful, in part, is because a human did it! If I am told a human did not do something, it almost immediately loses interest, not because of any inherent quality that is visible in the art, but because a human didn’t do it.

    For example, a hyper-realistic drawing is interesting not because it copies exactly what a photo does, but because it represents the efforts of an actual human. It symbolizes their toil and effort. It is similar to the reason we value a Bitcoin or any type of money: because it represents actual work. If something didn’t take work to do, we no longer value it.

    In other words, art is beautiful because it is proof of work, proof of human passion and insight and feeling.

    That being said, I think we also need to examine from first principles exactly why mathematical proof is beautiful, or interesting, or even worthwhile. I think there is actually an inherent human quality to mathematical proof that makes it beautiful.

    Moreover, the AI tools lack our inherent instinct for beauty. I don’t know how to describe what that is, and that’s the problem. I just simply have it.

    The other thing to note is that these AI tools have to use a lot of human inputs to work! So let’s suppose the human inputs stop. Will the AI continue to advance? I would guess not. Sure, they could keep producing various proofs. But without an instinct for recognising beauty, these proofs will not have relevance to humans. Indeed, almost all mathematically valid chains of reasoning are of no interest to us.

  61. Maybe An AI Says:

    2022: Pfft! AI can’t even do high school math!
    2023: Pfft! AI can’t even do college-level math!
    2024: Pfft! AI can’t even do PhD-level math!
    2025: Pfft! AI can’t even do math research!
    2026: Pfft! AI can’t even solve open research problems harder than the Erdős problems!
    2027: Pfft! AI can’t even win the Abel Prize!
    2028: Pfft! AI can’t even win a second Abel Prize!

  62. Heavy as a Charm Quark Says:

    > And that, whenever that happens, there will be new confident reasons not to care immediately offered up in comment sections like mine?

    It’s as sure as the sunrise over your sink.

  63. Lath Says:

    Thank you for adding the endnote. I wanted to write a similar comment but you phrased it way better than I could.

    Folks, it’s completely normal to be freaked out by this but stop moving the goalposts and have some intellectual honesty to admit the pace of progress has been staggering and shows no signs of slowing down.

  64. Richard Says:

    “Write me a story about the most ancient Israelites that’s riveting like the stories of the Bible but that’s also consistent with all of the archeological evidence.”

    Here is the story: Genetic evidence shows today’s Palestinians are the direct descendants of those ancient Israelites, while today’s Ashkenazi Jews are settlers from Eastern Europe.

  65. Not Owl Sowa Says:

    “Does anyone seriously doubt at this point that major open problems in algebraic geometry and other “Grothendieck-friendly” areas of math will fall to future AI models?”

    If such AI models are built, wake me up and I would like to ask them about how to prove or disprove Grothendieck’s standard conjectures on algebraic cycles. Algebraic geometers would learn a lot.

    But besides the fact that there is little evidence so far about AI changing the picture of algebraic geometry, there is also a huge gap between two metrics: 1. outperforming humans and 2. solving the Riemann hypothesis or the standard conjectures. Very likely AI models can do 1 very well (at least on Olympiad problems and combinatorics, as we have already seen), but doesn’t help with 2.

  66. Anon Says:

    #61

    anyone who has seen AlphaGo and knew a little bit about proof search could have seen that the answer to all these questions can be yes. as such there is little surprise in seeing these.

    however that is not the claim that is being challenged.

    has OpenAI punished the full transcript of what the model has done internally, including any sub-agents, failed exploration branches, calls to tools, …? until we have that, this is an impressive proof search system with an LLM heuristic, and the same way I would consider a very efficient non-LLM proof search system that finds new proofs very interesting and impressive and helpful but not very surprising, I would put this on the same category.

    we will soon have oss systems based on DeepSeek and other models that given enough compute would be able to find the same or similar things, and that is when we will start to have a more realistic understanding of what these systems are capable of and what they are not capable of.

  67. Anon Says:

    I am also waiting for the first AlphaGo move #37 in AI-assisted math results.

    but I would also remind that even though AlphaGo was able to defeat the top human player, many lower ranked players later on adopted and learned to defeat AlphaGo. that is what is not reported that an amateur player find a way to reliability defeat AlphaGo level AI players. why is that not as widely reported? because that does help the AI hypers.

    they always tell you only about positive results and what it can do and don’t bother mentioning what it fails to do.

  68. Mikhail Says:

    Prof.Aaronson, your posts about the LLM read like blogs of the people who have “relationships” with the silicone sex dolls. Could you please disclose whether you hold investment positions in “AI” stocks, have consulting contracts, grants or other business relationships with those companies?

  69. Balazs Says:

    I will reply to @Scott and @Darcyn at once.

    @Scott: you may call them blankfaces (and maybe rightfully), but the name-calling version of your argument is carbrain. That is seeing only two options (people driving lots of cars and killing people, vs. AI driving lots of cars and killing far less people). Helsinki in Finland achieved zero road fatalities last year, by implementing what we could call a different kind of urbanist innovation (not the VC backed kind, but one that outperforms AVs by every measure). People concerned about traffic deaths should maybe reflect on, why the political system in the US not able to deliver high-speed rail and dense tram networks, and much better at delivering VC funded big AV tech flooding the streets. (Even though the former clearly outperforms the best version of AVs in every human wellbeing measure – you can’t go lower than 0 fatalities, which is demonstrably achievable and only depends on political will.) Widespread AVs will inevitably further lock in governance and infrastructure in car-dependency. I realize that mine is a very European centric argument, and I would be fine with the US doing whatever if they feel they’re a lost cause for human centric urban development, but we’re getting AV tech now in the EU whether we want it or not, Waymo already lobbying against bike lanes in London.

    @Darcyn: the honest version of your argument is a simple question: how much should private companies pay for using the road infrastructure, congestion, and cannibalizing public transport? They must pay, otherwise you’ll get people making lots of low cost, but also low value trips and create lots of externalities (congestion, loss of even more public space to cars). I would be fine wit something like “cool tech, waymo, please pay $2 per passenger km for using public infrastructure and causing externalities”. But you don’t see such discussions as we are terrible with pricing traffic modalities (planes notably don’t pay fuel tax, and congestion pricing is very difficult to implement). So as long as AV pricing is not a prominent part of the discussion, I will oppose them anticipating an extremely unfair competition with other traffic modalities. Especially in cities with viable alternatives. (No wonder, that Waymo is pushing for at least nation-wide AV regulation in the EU, so that individual cities can not restrict or price them).

  70. anon Says:

    In my view, the problem is not that people are against generative AI, but rather that AI corporations are currently on the path to an IPO and are generating a lot of hype. In the same way that we don’t believe politicians when they are on the campaign trail, we cannot trust everything a corporation says or does when they are on an IPO trail. Therefore, as scientists, we must strive to keep the interaction between generative AI and mathematics strictly scientific. This implies that we should only accept generative AI capabilities for finding and proving new theorems through tests or experiments whose conditions we can fully control.
    This requirement excludes all results obtained via the internal models of AI corporations. Company-internal models are essentially black boxes: nobody, except the corporations themselves, knows how these results were obtained. Consequently, the public cannot control the conditions of these experiments, meaning their results cannot be accepted. For the same reason, I believe experiments conducted by individual scientists using released commercial models should also be excluded unless they have been performed under strictly controlled conditions.
    There have been initiatives to conduct controlled experiments. The one I am familiar with is named “First Proof.” This experiment was conducted in February of this year, and the results reported in their paper clearly show that the most advanced commercial AI models were not particularly successful in proving new theorems.

    Another experiment, which is under continuous execution, is the Erdős problems experiment. In this case, the testing is generally not done in a controlled manner, but we do have some statistics, at least from one of the corporations. I am referring to the Google Aletheia paper. Again, the results do not seem particularly favorable for generative AI. They tested the model against all seven hundred open Erdős problems and were truly successful in only two of them. Even though this specific corporation is not on an IPO trail, I personally do not fully accept these results because they were obtained using an internal model. Therefore, perhaps even these two results are not genuine. Naturally, I do not accept the claims that the Erdős unit distance problem was solved by generative AI.
    In summary, results obtained by AI corporations using their internal models cannot be accepted, as these experiments have not been conducted under controlled conditions. I would also extend this exclusion to results obtained by individual scientists in an uncontrolled manner; they simply cannot be accepted. We can only accept the kind of experiments that “First Proof” is conducting—specifically, the parts they control themselves. Personally, I think it is unfortunate that they called for community collaboration. While that represents the kind of popular science we all enjoy, it is strictly unscientific because you lose control over the experimental conditions.
    You, in this blog, have magnificently tried to distinguish hype from real progress regarding quantum computing experiments. I see no reason to treat generative AI any differently.

  71. Snurre Says:

    To quote Benedict Evans:

    People who are determined to believe that AI is fake move the goal posts almost as much as people who talk about AGI. First none of this works, then it works but it’s useless, then no one will pay for it, then they’re paying for it but not enough, then they’re paying $100bn this year, but it’s not profitable yet, then it solved a math problem but it’s not “real” math, and then…

  72. Adam Treat Says:

    I think double blind tests re: the Monet painting are an excellent idea. Make the people who wish to poo poo on AI actually test their criticisms in a double blind test.

  73. Adam Treat Says:

    Scott #34,

    “(2) getting past regulatory blankfaces who are congenitally unable to think statistically (“we’re already getting into 90% fewer crashes than human drivers!” “AHA, then you admit you’re not 100% safe!”).”

    I think this is the case in coding as well. Someone upstream was noting that productivity gains haven’t materialized as much as you’d expect considering how good the AIs are at coding. I suspect that the problem is that *human* software engineers are standing in as “regulatory blankfaces” much more than they need be.

  74. Andy Pezzi Says:

    I don’t think AI will ever be able to reframe the mathematical “scaffolding” of quantum mechanics on the base of what is, for my money, the most “quantum” aspect of quantum theory, that is: entanglement – or the non-factorizability of the quantum systems fully “quantum”.

  75. Scott Says:

    Richard #64:

      Here is the story: Genetic evidence shows today’s Palestinians are the direct descendants of those ancient Israelites, while today’s Ashkenazi Jews are settlers from Eastern Europe.

    Thank you for illustrating, better than I could, the loony, debunkable Jew-hating conspiracy theories at the core of so much antizionism!

    Genetic evidence confirms the obvious thing that no one even thought to deny until quite recently: namely that modern Palestinians, and Ashkenazi Jews, and Sephardic and Mizrahi Jews (as numerous as Ashkenazim in modern Israel, after the Arab countries where they’d lived for centuries expelled them), all have genetics that trace back to the Levant region, where the ancient kingdoms of Israel and Judah that gave the world the Bible once stood. In all three cases, intermixed with the various other places where their ancestors lived.

    If this weren’t the case, and even if we didn’t have the DNA evidence, what exactly is the alternative theory for how those … uhh … “Eastern European settlers” retained what anyone can see are a language and religion from ancient Israel? Why was return to Israel (“next year in Jerusalem”) at the core of their liturgy for 2000 years of exile? Are they all Khazarian imposters who just decided to adopt this ancient culture as a skinsuit? (DNA evidence has refuted that particular speculation.)

    The funniest part is that even Jew-haters like yourself used to know all this stuff. Before most of the Jews of Europe were outright murdered in the Holocaust, they were often told to “go back to Palestine.” But antisemitism is a notoriously shape-shifting mental illness, one that adapts its wild conspiracy theories to whatever are the current circumstances of the world and its reigning ideological frameworks, from Communism to nationalism to Islam and Christianity to “settler colonialism,” even while it never varies its target.

  76. JimV Says:

    Let me be the first to say: great post! (I’ve read the 70 comments shown at the current time and if it has already been said, I missed it.)

    My position has been that the recent (maybe not current) AI era has been like the Model-T era of automobiles (“Get a horse!”). AlphaGo convinced me that we have a self-learning computer algorithm-system. (Yes, I am aware there was a strategy found which was not in AlphaGo’s training set, but which the next version of AlphaGo, trained by playing itself, is not susceptible to.)

    At the same time, I agree that most of the “AI Assistants” at commercial websites are currently mostly ineffective and somewhat annoying.

  77. Mitchell Porter Says:

    This is just a microcosm of what is now happening to human intellect in general. I am old enough to be instinctively wary of proclaiming the end of the world, but it does look like the human era on Earth is ending. Whether humans will literally go extinct because we can’t keep up with whatever new imperatives govern the world, or whether a safe space will be kept available for us, are now serious questions about the posthuman order.

  78. Joshua Zelinsky Says:

    Anon #70:

    This requirement excludes all results obtained via the internal models of AI corporations. Company-internal models are essentially black boxes: nobody, except the corporations themselves, knows how these results were obtained. Consequently, the public cannot control the conditions of these experiments, meaning their results cannot be accepted

    People have used proprietary software before AI. Someone can for example use Mathematica to factor some very large number and then verify that factorization. I don’t need access to the Mathematica sourcecode to check this, and no one has any doubts that Mathematica really did factor it. But for someone reason, with LLM AI, a much higher bar is insisted upon.

    Naturally, I do not accept the claims that the Erdős unit distance problem was solved by generative AI.

    So instead it was what? What is your alternate hypothesis? That mathematicians worked with OpenAI to solve an extremely difficult problem, came up with a solution which was opposite the direction anyone else expected, and then deliberately decided to hide they had done that, and then instead produced an 125 page fake output log? Do you see why when that’s stated explicitly that has far more difficulty than the alternative?

  79. Richard Says:

    @Scott 75: Some Mizrahi Jews yes, but Eastern European Ashkenazi Jews clearly no.

    “what exactly is the alternative theory for how those … uhh … “Eastern European settlers” retained what anyone can see are a language and religion from ancient Israel?”

    uhh… they spoke Yiddish… not exactly a Semitic language… and they are converts to Judaism… just like black Jews and many Sephardic Jews (and Marilyn Monroe).

    The genetic evidence is clear: Palestinians and Samaritans are the descendants of ancient Israelites while Ashkenazim are Eastern European converts and modern settlers…

    Sorry I had to break it to you… but facts are facts. Even some Israeli newspapers have acknowledged it. Did you really think those guys from Lithuania and Belarus were Semites…?

  80. Scott Says:

    Joshua Zelinsky #78:

      What is your alternate hypothesis? That mathematicians worked with OpenAI to solve an extremely difficult problem, came up with a solution which was opposite the direction anyone else expected, and then deliberately decided to hide they had done that, and then instead produced an 125 page fake output log?

    Beautifully put! Whether with antisemitism or vaccines or quantum mechanics or AI, laughable conspiracy theories are sort of the ultimate diagnostic for when one side has no remaining arguments but also no ability to admit it.

    Imagine if we applied anything like commenter #70’s newly invented standard to human mathematicians. Imagine that, after Wiles announced his proof of Fermat’s Last Theorem, people said to him: “see, the problem is that your brain is a black box, whose internals the rest of the world doesn’t know. Thus, until surgeons can dissect your brain on an operating table, and understand the results well enough to piece together every step of your research journey—naturally, verbal reports won’t suffice—we can’t accept that you actually did prove the theorem you claim to have proved. Sure, the proof itself looks good, but who knows where it may have come from? Maybe God or aliens just handed it to you!”

  81. Scott Says:

    Not Owl Sowa #65: Very well then. I take you to have registered a prediction here, that AI will not solve any significant open problems in algebraic geometry within (say) the next few years.

    I, too, wish to register a prediction: that AI will solve significant open problems in algebraic geometry—and that when it does so, you’ll be here ready with arguments that the problems in question weren’t the truly most fundamental ones, and in any case the fact that AI solved them isn’t particularly surprising or interesting and shouldn’t count.

  82. Scott Says:

    Anon #67:

      but I would also remind that even though AlphaGo was able to defeat the top human player, many lower ranked players later on adopted and learned to defeat AlphaGo. that is what is not reported that an amateur player find a way to reliability defeat AlphaGo level AI players. why is that not as widely reported? because that does help the AI hypers.

      they always tell you only about positive results and what it can do and don’t bother mentioning what it fails to do.

    And how are the top human Go players doing right now against the top Go AIs? Come on, since you’re so insistent on the complete story being told: how are they doing?

  83. kamil Says:

    I, for one, welcome our robot overlords. Living in a world ruled by Claude would be probably much better for the median person than what we have now. While we should not downplay the risks on this path, ultimately AI augmenting and replacing humans is a desirable outcome. Let’s just make sure the Claudes win instead of sociopathic AIs.

    @Balazs:
    As a fellow European, I find this argument against self-driving cars very weak. By and large these two issues are totally orthogonal. Both reducing traffic AND allowing self-driving cars, buses, trains, lorries, etc. will help reduce traffic fatalities. Sure, self-driving cars may introduce perverse incentives here and there but this won’t be enough to offset the benefits. Regardless, it would be very naive to think that the US will reform their urban planning to become the next Helsinki, so self-driving cars are the next best thing, and they would make Helsinki, too, even more awesome.

  84. Scott Says:

    Richard #79: Yes, they spoke Yiddish … which is written with Hebrew letters! While they also maintained Hebrew as a liturgical language that they all needed to keep learning, well enough that they then successfully revived Hebrew as a daily use language in the late 19th century, and well enough that modern Israelis can understand the Bible as originally written at least as well as modern English speakers can understand Shakespeare.

    Of course there were always converts to Judaism. According to DNA analysis, modern Ashkenazim seem to be largely descended from Levantine men who intermarried with Italian women after the Romans’ destruction of the Second Temple. Here, let’s ask Google AI about this:

      Yes, modern Ashkenazi Jews have significant Levantine DNA. Extensive genetic studies show that they possess a robust Middle Eastern genetic core, typically tracing about 50% to 60% of their ancestry back to the ancient Levant—alongside European (predominantly Southern European) admixture accumulated over centuries.

    Or let’s ask Wikipedia, which despite having been taken over by anti-Israel fanatics in the last decade, hasn’t yet been able to scrub this particular reality from the record:

      Like other Jewish ethnic groups, the Ashkenazim originate from the Israelites and Hebrews of historical Israel and Judah. Ashkenazi Jews share a significant amount of ancestry with other Jewish populations and derive their ancestry mostly from populations in the Levant and Southern Europe.

    Not that I expect that citing even a thousand such sources would penetrate your mental illness.

    Indeed, I notice that you don’t even bother to deny my description of you as a “Jew-hater.” So then, Jew-hater: did you really imagine that I knew my own people’s history so poorly that you could gaslight me about it?

  85. Joshua Zelinsky Says:

    Richard #79.

    So, everything you’ve written is wrong or deeply misleading. Let’s go sentence by sentence.


    Some Mizrahi Jews yes, but Eastern European Ashkenazi Jews clearly no.

    So, aside from this claim being wrong, over half of Israelis Jews have Sephardi ancestry https://en.wikipedia.org/wiki/Israeli_Jews


    uhh… they spoke Yiddish… not exactly a Semitic language… and they are converts to Judaism… just like black Jews and many Sephardic Jews (and Marilyn Monroe).

    Yiddish is a highly semitic language. It is a mix of German, Hebrew and Aramaic. https://en.wikipedia.org/wiki/Yiddish . This is a pretty basic fact. What’s the next step? Claiming that Mizrachi Jews who speak Ladino aren’t really descended from the area because it contains a lot of Spanish and Latin elements?


    The genetic evidence is clear: Palestinians and Samaritans are the descendants of ancient Israelites while Ashkenazim are Eastern European converts and modern settlers…

    Sorry I had to break it to you… but facts are facts. Even some Israeli newspapers have acknowledged it. Did you really think those guys from Lithuania and Belarus were Semites…?

    Making claims like this without even trying to back it up is impressive. It is even more impressive for the high condescension of the writing style along with being just *wrong*. First, all major Jewish populations share a lot of genetic traits. See https://www.science.org/content/article/tracing-roots-jewishness and cell.com/AJHG/abstract/S0002-9297%2810%2900246-6# . While it is true that there is some European ancestry of the Askhenazic population it has a large amount of Middle Eastern elements. See https://pmc.ncbi.nlm.nih.gov/articles/PMC9793425/ https://pmc.ncbi.nlm.nih.gov/articles/PMC5380316/ . There some evidence that there’s more matrilinean European elements and more patrilineal Levantine elements, but that’s not relevant here. If you want a summary with less of the technical detail and a really good visual diagram see https://www.razibkhan.com/p/rkul-hit-list-2023-ashkenazi-jewish .

    And then, when you’ve read all of that, I suggest you take a look in the mirror and ask what made you uncritically accept this claim. It doesn’t make you an antisemite, but it does suggest you are willing to listen and take for granted claims being made by people who are not trying hard to make their beliefs match the evidence we actually have.

  86. Scott Says:

    Mikhail #68:

      Prof.Aaronson, your posts about the LLM read like blogs of the people who have “relationships” with the silicone sex dolls. Could you please disclose whether you hold investment positions in “AI” stocks, have consulting contracts, grants or other business relationships with those companies?

    Very well: at this moment, I have no investments in AI companies, and no consulting contracts or grants or business relationships with them. I certainly might in the future, and of course a few years ago I did work for OpenAI—a company many of whose recent choices I’ve loudly criticized since I left.

    I personally know quite a few people who’ve become paper billionaires from their involvement with AI companies, but I have not. Instead I’ve just remained a professor, with a professor’s salary, trying to tell the truth about everything, and desperately arguing with ideologues like yourself who denounce me on the Internet—even while my friends just laugh at the ideologues, ignore them, and get rich. No doubt my friends are much wiser than me in this respect.

  87. Hyman Rosen Says:

    I was referring to this: https://www.science.org/doi/10.1126/science.adx2728
    It’s all very recent. Sorry I don’t know enough about QC to say anything intelligent about it. Note that unlike with respect to AI alignment, I am not in general a skeptic about quantum computing. If we achieve actual quantum supremacy with actual quantum computers, that would be great, although it would also be neat if the universe somehow conspired against that being possible. I’m reminded of quark containment, where the universe doesn’t let you isolate quarks by trying to split apart particles made of them – the energy just produces more sets of contained quarks.

  88. MD Says:

    JimV #76: Can you please share a link to this next version of AlphaGo?

    The last that I have seen on this topic was a paper from 2024 (https://arxiv.org/abs/2406.12843) which showed that the adversarial attack still works on KataGo after a few attempts to patch it. But KataGo is not AlphaGo, so you might still be right…

    For what it’s worth, I tried asking various free-tier LLMs and couldn’t find what you’re referring to, since they tended to point me to either the 2017 AlphaGo paper or the 2023 “Adversarial Policies Beat Superhuman Go AIs” paper, which are both not what I’m asking for.

  89. Ryan Says:

    Hey Scott,

    I can’t believe this shit. Are you serious? You used to be a smart and skeptical guy, but you read the headline and just swallowed it?

    Read Marcus.

    https://garymarcus.substack.com/p/checking-the-math-behind-openai-and

    The whole thing is so much thinner than the headline. “AI solved an Erdős problem.” No, what happened is that an unreleased model dumped out a long fucking transcript and actual mathematicians went through it with tweezers.

    Marcus quotes Cal Newport:

    “Professional mathematicians identified the counterexample from within a long transcript of the model’s reasoning, and then extracted the key parts and rewrote it as a more succinct proof in a more standard style.”

    I mean, come on. That sentence tells all. They cherry-picked one accidental success from a massive pile of shit.

    The machine did not hand in a proof, it did not know it had a proof, it did not know which thoughts were useful. It output a pile of garbage. Actual mathematicians searched the pile. Then actual people wrote the proof.

    This is being sold as if the machine did all of this! It’s a fraud.

    Also nobody tells us how big the pile was. Marcus again:

    “We also don’t know how many prompts they tried, and how many didn’t work; we have a numerator but not a denominator.”

    That is the part everyone wants to skip because it ruins this AI is awesome tech bro story. Maybe it took a few tries. Maybe it took a stupid number of tries. Maybe most of the transcript was wrong. Maybe most of the runs were worthless. We do not know. They know, but they aren’t telling us.

    And the model is unreleased, naturally. So you can’t reproduce it, you can’t inspect the prompts. You just get the fucking oress release from OpenAI.

    Marcus says we do not know how the model worked how it was trained, or how general the result is. He also says there is no data about whether this tells us anything about hallucinations or other benchmarks, or cost.

    So what exactly are people celebrating? A real counterexample, maybe. Good. Fine. But the AI techbro mythology piled on top of it is utter garbage.

    And the Erdős name is doing half the work. People hear his name and think “if he couldn’t do it then only a genius could solve it.” But this was a counterexample hunt in discrete geometry. That is exactly where you might get value from generating a bumch of weird candidates and then checking them. That is not the same as understanding a theorem. That is not the same as proving a deep result in algebraic geometry for example, something that requires actual foresight and theorem building instead of generating a bunch of possible counterexamples and then having human mathematicians check them.

    This is no better than “computers proving the four color theorem” which happened in the 70s, and you are a fool for thinking otherwise.

    Thomas Bloom, quoted in the piece, says the model had “superhuman levels of patience.” Yes. That sounds right. Patience is cheap when you are a machine being paid for by someone else’s compute budget and by all the fresh water and electricity and carbon in america. Patience is not taste or judgment or evenknowing what the fuck you are doing.

    Newport says, also quoted by Marcus:

    “I don’t think it’s accurate to say these examples of AI-supported mathematics mean the models are somehow ‘smarter’ than human mathematicians.”

    Good. Then say that. Maybe AI can be a candidate generator for some counterexamples. More or less by stochastic stumbling.

    But “AI solved an Erdős problem” is not a sane claim. It’s fucking marketing for OpenAI.

    Newport even says the experiment might be “more about marketing the power of their new model than trying to actually advance computer-aided math.”

    Show the raw transcript and all the failed attempts and then maybe we can talk about this..

    Until then, the accurate version is boring:

    An unreleased OpenAI model produced a long transcript. Human mathematicians found a counterexample inside it and rewrote the proof.

    It’s a fucking Clever Hans. It’s monkeys punching a keyboard. In other words, a stochastic parrot. And honestly, shame on you for being snookered into their marketing scheme.

  90. Scott Says:

    kamil #83: Indeed, it’s noteworthy how opponents of self-driving cars are forced to change the subject to something orthogonal.

    I, too, wish we had reliable high-speed trains all over the US. That also happens to be an issue of regulatory blankfaces—the ones who’ve made new construction essentially impossible, by tying it up forever in lawsuits and environmental and labor reviews. The blankfaces ironically set back the very causes (like affordability and the environment) that they claim to support, once again because they’re unable to think globally and quantitatively or don’t want to. See Klein and Thompson’s Abundance book for much more about this.

    Yet none of that changes the fact that full self-driving would save 40,000 lives per year in the US alone, with our existing infrastructure. To anyone looking at the thing on the merits, rather than as an ax-grinding ideologue, it would be a no-brainer, and would not be either/or.

  91. Scott Says:

    TO EVERYONE SANE:

    If you’ve read the comments down to this point, then you now understand one of the main reasons why I blog so much less than I used to!

    When I travel to give lectures in real life, person after person tells me how much they love this blog, how they’ve been reading it since high school and how it inspired them to go into science. They want me to sign their copies of Quantum Computing Since Democritus, and to take selfies with me. They ask reasonable questions about quantum mechanics and complexity theory and AI, and care what answers I give. They teach me things I didn’t know. Sometimes they want to hire me as a consultant or for their company’s board.

    But then in this comment section (and on Twitter and Reddit), and only there, I’m a bloodthirsty genocidal Zionist, and also a cynical shill for the AI companies, and also a naïve fanboy for them, and also such a laughably bad mathematician that I imagine solving an Erdös problem to be impressive. What an utter piece of shit I am!

    And yet, it’s almost impossible to tell which of the endless pseudonymous attacks are serious and which are trolling, which from humans and which from AI. Even the ones from humans seem like they might as well be from particularly simple chatbots, ones that never update in response to anything.

    So then, whenever I’m wavering between putting up a new post or not, it’s like: why not spare myself the grief?

  92. Scott Says:

    Ryan #89: See, but that’s not how Timothy Gowers or Noga Alon or Will Sawin or Daniel Litt characterized the situation, is it?

    The actual experts all said: GPT supplied the essential idea of the counterexample, in response to a prompt simply asking it to solve the unit distance problem. And in this case, the answer was not buried in a gigantic pile of shit (although that would’ve been interesting as well): once GPT had the correct idea, it “knew” it had it, and it said as much in its chain-of-thought transcript. The humans’ role was mostly to understand and explicate the idea.

    So, are Gowers and Alon and all these other world-renowned mathematicians just shills for the AI companies? How incredibly far-reaching are the tentacles of this conspiracy! Maybe you, too, are just a shill, being paid to illustrate the feebleness of the opposing side!

    As for Gary Marcus, I do have some common ground with him—he and I even joined forces to speak out against some of OpenAI’s actions over the past couple years. But yes, Gary has backed himself into a corner where he’ll constantly need to fight a rearguard action, will constantly be surprised by the next thing AI accomplishes, and will then seek whatever rhetorical ways he can to downplay and minimize that thing.

  93. Ryan Says:

    Scott,

    This “argument by authority” attack doesn’t work. Daniel Litt is extremely woke, politically correct, and anti-israel per his “bluesky” page, where he even has pronouns listed, so why should we trust anything he says? Timothy Gowers is perhaps skilled in CS-adjacent discrete math but not a real serious bona fide mathematician. There’s a blog “stop timothy gowers!” by a respected algebraic geometer you should read. Modern algebraic geometry is leagues above combinatorics and graph theory, in abstraction and difficulty..in Gowers’ work there is essentially no theory-building, just explicit examples and combinatorial reasoning. of course computers can do that without being geniuses

    And yes I am very upset about this. This AI bullshit is threatening my career in math. What if universities won’t hire mathematicians anymore because they’ll have AI teach all the classes, and publicity stunts like this convince them that AI will do research too? What if all the math jobs go to this ML AI bullshit? What’s left for real mathematicians like me? And on top of that, the AI is destroying our water and our environment with data centers going up everywhere (and unlike with most environmental bullshit the Left makes up, this is real this time), and it’s making everyone stupid and everyone is depending on it. Of course I am angry at the enablers.

  94. LK2 Says:

    I hope you did not include my post in the “people willing to diminish the achievements of AI”.

    I just asked if discrete math fits better the present models and if there are any achievements in “continuous math”. In summary, I wanted to understand the capabilities of present models, without implying that in the future AI will not reach the creativity of Groethendick et similia. Might well be.
    Personally, I’m astonished by their capabilities already now and use LLMs every day for academic research (physics).

  95. Daniel Litt Says:

    Ryan #89:

    You quote Cal Newport saying:

    “Professional mathematicians identified the counterexample from within a long transcript of the model’s reasoning, and then extracted the key parts and rewrote it as a more succinct proof in a more standard style.”

    This is false. As I understand it, the model accurately identified that it had a proof and produced a TeX writeup. I and the other authors of the commentary paper checked, digested, and simplified the argument a bit–the models still leave much to be desired as mathematical writers–but the model produced a perfectly understandable writeup on its own.

    The accomplishment here is very real. There are still plenty of mathematical tasks that the models are not yet particularly good at, but obviously the trajectory is remarkable and it doesn’t make sense to deny this.

  96. Dacyn Says:

    Balazs #69: OK, that is a more reasonable argument than what you wrote in #50. I don’t know that I agree but I’m happy to end the conversation here.

  97. Mike-e Says:

    The coping problem isn’t with “the achievements of AI” per say, but that those achievements really underline that the human brain isn’t that special after all. That something as simple as an LLM can match it easily (many competing products are on par), and that it’s a matter of quantity (of parameters matching neuron count) over quality (no super intricate algorithm needed).
    And the fact that cognition and consciousness are two separate things.

  98. Daniel Litt Says:

    Ryan #93:

    For what it’s worth I don’t think your description of my political views is accurate, nor is it relevant to my ability to judge the quality of the work here. In any case, what you’ve written strikes me as a remarkable example of motivated reasoning. If you are indeed a mathematician you should be able to read the counterexample to the unit distance problem and make a judgment for yourself.

  99. Adam Treat Says:

    “And yes I am very upset about this. This AI bullshit is threatening my career in math.” — Ryan #93

    “It is difficult to get a man to understand something, when his salary depends on his not understanding it.” — Upton Sinclair

  100. Adam Treat Says:

    Ryan #93,

    You’re saying two contradictory things:

    – The AI did nothing good here! It was actually the human mathematicians who did all the work and the AI wasn’t even responsible!

    – The humans who did all the work are actually crappy mathematicians so we can’t trust them when they say that the AI did any of the work!

    Perhaps reflect on this. The logical error here is so basic I do indeed wonder about your long term employment prospects in math.

  101. Mike-e Says:

    I think AI performance will really explode once it can break free from the limitations of human language (a mere 60,000 words or so) and starts augmenting it with millions of new concepts- like how Eskimos have a dozen words for all the subtle variations of snow, but for a dude from Texas it’s all just “snow”.

  102. Scott Says:

    I’ve reached the end of my amusement with the troll “Ryan,” and am leaving his further comments in moderation.

    While I take a backseat to no one in academic STEM in my public Zionism, I’m also perfectly able to decouple. There are many, many antizionists and others who probably hate my politics, who I nevertheless trust in their areas of scientific expertise more than I trust myself. As for Daniel Litt’s views, I never thought to check and have no idea.

    Daniel: Thanks so much for your efforts to explain the progress on the unit distance problem, here and elsewhere!

  103. Joshua Zelinsky Says:

    Ryan #89 and Ryan #93,

    Since both of your comments have a lot of overlap, I’m going to comment on them together. Since some of the issues have already been pointed out, I’m going to try and only address aspects which have not been addressed yet.


    I can’t believe this shit. Are you serious? You used to be a smart and skeptical guy, but you read the headline and just swallowed it?

    Generally, if someone is someone you think is “smart and skeptical” and they disagree with you on something you think is obvious, that should cause you to update in the direction that your belief that the thing is obvious or clear cut is not as valid as you think otherwise.


    No, what happened is that an unreleased model dumped out a long fucking transcript and actual mathematicians went through it with tweezers.

    It is important to recognize here that the AI said essentially “yeah, I’ve got a proof.” Even if the AI said “there’s a proof somewhere in here,” that should strike you as impressive.


    “Professional mathematicians identified the counterexample from within a long transcript of the model’s reasoning, and then extracted the key parts and rewrote it as a more succinct proof in a more standard style.”

    I mean, come on. That sentence tells all. They cherry-picked one accidental success from a massive pile of shit.

    You claim to be a mathematician. So I’m going to ask if my workflow is similar to your workflow. My own workflow on a problem often goes something like this when I work on a problem. I try one thing. It fails. I try to look at a related theorem; doesn’t generalize. I go check a few cases; doesn’t give much insight except for a pattern which breaks down soon, so I rope an undergrad or high school student into writing some code to check a lot of examples. They learn something, and we notice a few patterns in the data, but nothing that useful or non-trivial. Then, I’m sitting in a seminar on a completely different topic, and trying to pay attention while the speaker does a really poor job explaining their research, I’m like “Hmm, what if I tried to combine it with that other thing I saw 2 years ago.” That still doesn’t work. But then six months later, I bash my head against the problem a bit more trying to use some sophisticated representation theory results, which fails (whether because it cannot work or because I just don’t know representation theory well) and then I’m falling asleep and I realize that other thing from now 2.5 years ago combines with a pattern the undergrad mentioned in the data that I didn’t think was important, and can get it to work with a linear algebra trick a student mentioned a few days ago. Does that sound familiar?

    The point I’m making is that an 125 page giant transcript most of which is junk but has a result embedded in it sounds a lot like what it would look like if one transcribed my own thoughts on a problem from start to finish, and I suspect that that’s pretty close to the case
    for many if not most other mathematicians.

    As for writing it up well: people clean up results from others all the time. I wrote two papers improving on a specific inequality due to Pascal Ochem and Michael Rao, and my second paper was a giant complicated proof, which someone else described as a “tour de force but not in a good way.” Then two other people looked really closely at all three papers and wrote a paper which improved the inequality even further, and in the process drastically simplified and cleaned up the whole method to the point where reading their paper made me feel like I had a much better understanding of what my own earlier papers were doing.

    I’m not the most skilled mathematician out there. My advisor called himself a “second rate mathematician” and I’m at least one rate if not two rates worse than him. So it is possible that all of these experiences are just the things that people who are the weakest mathematicians experience. But I doubt that.


    Also nobody tells us how big the pile was. Marcus again:

    “We also don’t know how many prompts they tried, and how many didn’t work; we have a numerator but not a denominator.”

    A denominator is interesting here, and it would be useful information. It is especially the case because they likely ran the AI against many different famous or moderately famous problems. But that still feels a lot like an actual mathematician. I’d guess that I make publishable progress on about a thirtieth of the problems I seriously think about, and that might be generous. So even if the denominator here is gigantic, it is acting a lot like an actual mathematician. Again, a weak mathematician like myself, maybe.


    That is the part everyone wants to skip because it ruins this AI is awesome tech bro story. Maybe it took a few tries. Maybe it took a stupid number of tries. Maybe most of the transcript was wrong. Maybe most of the runs were worthless.

    I’m not sure how that ruins anything. If the AI has to run through a 1000 runs but it was able to make a run that worked and told the humans that time “Hey! I got a result!” That’s great. If you could copy a mathematician into a thousand separate parallel universes with slightly different starting points and then select at the end the one which actually succeeded on a problem, that would be incredibly powerful. How does doing massive runs diminish it?

    And the Erdős name is doing half the work. People hear his name and think “if he couldn’t do it then only a genius could solve it.” But this was a counterexample hunt in discrete geometry. .

    You can make that argument if you want as a problem for some of the weaker Erdos conjectures. And there’s a problem in general with the idea that one needs to be a “genius” to do good math, although it does certainly help. But we’re not talking about a random Erdos problem; we’re talking about a very well known problem where a lot of people had thought about it.


    That is exactly where you might get value from generating a bumch of weird candidates and then checking them. That is not the same as understanding a theorem. That is not the same as proving a deep result in algebraic geometry for example, something that requires actual foresight and theorem building instead of generating a bunch of possible counterexamples and then having human mathematicians check them.

    Have you read the paper in question? The construction isn’t some specific giant example, like “Hey, here’s a specific set of points where it is large.” . It is a construction using a tower of number fields. It is just as much a theorem as anything else.

    That said, t may be that discrete geometry and combinatorics has an easier time with these sort of problems. I suspect that a lot of the Erdos problems have turned out to be vulnerable in part because they require less technical background, so easier for both humans and AI to understand, and because they are reasonably simple, they are more prone to people just throwing them at the AI because one doesn’t need to have deep knowledge to give it to the AI.


    Thomas Bloom, quoted in the piece, says the model had “superhuman levels of patience.” Yes. That sounds right. Patience is cheap when you are a machine being paid for by someone else’s compute budget and by all the fresh water and electricity and carbon in america.

    And how much water and electricity and carbon would it have taken to generate a similar result with that many mathematicians? Incidentally, while the electricity used by data centers is a lot, any given run is pretty small. And water usage is drastically overestimated. See e.g https://theconversation.com/ai-has-a-hidden-water-cost-heres-how-to-calculate-yours-263252 here for a start. (Also the push to reduce water use is causing data centers to switch to more energy intensive cooling systems so this is also actively unhelpful if you are concerned about the carbon production and electricity use.)


    It’s a fucking Clever Hans. It’s monkeys punching a keyboard. In other words, a stochastic parrot.

    Aside from the apparent weird anger in this, and all the other problems already outlined with this, and ignoring how a mere “stochastic parrot” would be able to tell the mathematicians it had succeeded, I want to point out that the “stochastic parrot” term is not just wrong about the AIs but also shows a lot of misunderstanding about parrot cognition and linguistics. Parrots, especially African Greys, but others as well, have demonstrated grammar as well as understanding large vocabs and broader context. The term is not just a bad one for AI but reflects a failure to appreciate animal cognition.


    This “argument by authority” attack doesn’t work. Daniel Litt is extremely woke, politically correct, and anti-israel per his “bluesky” page, where he even has pronouns listed, so why should we trust anything he says?

    So aside from the problems of a general ad hominem, it is worth recognizing that you can disagree with someone, even disagree with them on a whole bunch of issues, and still listen to them where they are an actual subject matter expert. I disagree with Litt on some things, and agree on others, but if he makes a point about math and AI, I’m going to listen since he’s one of the people who has tried to work with these things since even before ChatGPT. It is also curious that you complain about Litt’s politics, but a lot of the things you’ve referenced are coming from the “woke” end of things. Labeling AIs as “stochastic parrots” for example came from Emily Bender and Timmit Gebru. Concern about AI water use is almost exclusively a left-wing, “woke” concern which often overlaps with anti-Israel views and that whole cluster.


    Timothy Gowers is perhaps skilled in CS-adjacent discrete math but not a real serious bona fide mathematician.

    Gowers is a Fields Medalist. In general, when one needs to discount the a Fields Medalist as not a serious mathematician to make your point, you really should wonder if something has gone drastically wrong. You shouldn’t take what either Litt or Gowers has to say for granted. But you should realize that they might be right; that these are people where mere dismissal of their ideas is not enough. They aren’t infallible; I helped answer a question Gowers asked once, albeit my contribution was an extremely trivial observation. But they are not just random individuals.


    And yes I am very upset about this. This AI bullshit is threatening my career in math. What if universities won’t hire mathematicians anymore because they’ll have AI teach all the classes, and publicity stunts like this convince them that AI will do research too? What if all the math jobs go to this ML AI bullshit? What’s left for real mathematicians like me?

    So here’s the thing: one should be able to separate what one wants to happen in the world with what is likely to happen. That can be tough. Suppose a giant asteroid might be hurtling towards Earth. Then figuring out if it is actually doing so is really important. Deciding that because one doesn’t like the consequences of that to get angry at the people who are saying it might hit Earth is not helpful. I would also suggest that taking a view that other people who are proving difficult results are not “real mathematicians” is not helpful. If you want to argue that people like me who have only minor results aren’t real mathematicians, then maybe feel free. But doing so based on not caring for their field of research is really unproductive.

  104. Alex Xanthakis Says:

    I think the “Monet experiment” however interesting says more about “art criticism” and “art critics” than GenAI and its future in general. The Erdős problems’ proofs and refutations by GenAI are seminal, and whoever belittles such achievements is hiding his head in the sand. Taking GenAI behaviour into consideration does this mean that in the near future Mathematicians will have to painstakingly go through GenAI proofs, searching for hallucinations, essentially ceding what makes them human while at the same time makinga handful of companies big $$$? Your post is a canary in the AI coal mine for young Mathematicians, and I expect a massive existential crisis for people working in the field. Cheers!

  105. Mike-e Says:

    The points of that “Ryan” character are so cleared contradictory bait that he must be a troll, haha.

    ”What if universities won’t hire mathematicians anymore because they’ll have AI teach all the classes”

    If that’s the case, there won’t be any need for math classes. Why would we train humans for an activity that would be better done by an AI?

    ”And on top of that, the AI is destroying our water and our environment”

    He’s the one sounding like a politically correct libtard there :-p

    Anyone whose current job consists in assembling long sequence of tokens in the right order should start looking for a new job…

  106. Christopher Says:

    Re: Scott 91

    Idea for a new policy: have a LLM that detects dumb questions automatically. If the question is dumb, the LLM automatically responds, saving you time and sanity! A LLM can give good faith answers to bad faith questions all day long.

  107. Scott Says:

    Hyman Rosen #87: Ah yes, I saw that paper. It looks like they managed to dequantize some condensed-matter simulation that was done on a D-Wave machine. Good for them, if perhaps unsurprising to anyone who remembers the days when this blog was Ground Zero for D-Wave skepticism! 🙂 There’s no claim to dequantize the simulations from Google, Quantinuum, or QuEra that are actually state-of-the-art right now.

  108. Christopher Says:

    The biggest question in my mind is when LLMs will start making breakthroughs in the following fields:

    1. Robotics research
    2. Generative AI itself

    Since that is when the physical and intellectual RSI will begin.

    I think these are a bit harder than disproving a single conjecture in geometry. The biggest gap is that LLMs are hugely biased towards verifiable problems, as noted by METR. There is no Lean checkable proofs in robotics or ML!

    But only a *bit* harder, I’m fairly confident that just scaling the current architecture will suffice even if it takes a few years. We’re still riding the exponential. And FigureAI might crack robotics soon anyways, no RSI needed.

  109. Mike-e Says:

    Alex

    “Taking GenAI behaviour into consideration does this mean that in the near future Mathematicians will have to painstakingly go through GenAI proofs, searching for hallucinations”

    another bunch of AIs will do the “peer reviewing”.
    It’s ok because humans won’t even be able to keep up anyway (humans can already keep up when the humans were writing the papers..).
    The only time humans will take note of some math/physics breakthrough is when it will have an engineering impact.

  110. Mike-e Says:

    It’s baffling that so many don’t realize that once AI is capable of doing “research” on its own, it’s GAME OVER for academia.
    You expect that by some magical hidden rule of fairness, AIs will just produce stuff at a rate on par with what humans in academia can produce or even just absorb?
    Even if qualitatively it would somehow stay on par with what a human can produce, it can be made faster and faster… producing a good paper in a second rather than a year.
    Remember Kurzweil and his SINGULARITY? Turns out he was right!

  111. Balazs Says:

    Scott #90: AVs vs other modes of transport are not orthogonal, as these compete for funds, and quite literally for space. My argument is that scaling up AVs should not only be the function of safety, but also that the playing field is level (which must involve some congestion/road pricing for AVs and possibly other restrictions). I understand why someone might object to this on moral grounds. “We can save lives now! Let’s sort out road pricing later!”. Except the later never comes, and the moral argument is in practice used to bypass the tougher governance questions that determine whether this technology makes cities better or worse in the long run. So @kamil no, we can’t just expect AVs to make cities better by default, the “perverse incentives here and there” is not an edge case, because unless AVs are priced for road use properly (and/or severely restricted), no other mode of transport can compete with them (not because they’re so great, they’re just cars, but because they free-ride the infrastructure the externalize costs, so bizarrely the least efficient mode of transport becomes the cheapest). Btw I also think human drivers should be subject to the same, so in that sense I am not anti-AV, only anti-car.
    I really do understand why insisting on any other condition than safety may look like I’m “an axe-grinding ideologue”. But consider that even for a medicine that works, the bar for allowing a drug on the market isn’t only that it passes clinical trials, a drug that works but is uneconomical (eg. so expensive that it takes money away from other more effective treatments) is facing regulatory roadblocks and rightly so.

  112. JimV Says:

    Reply to MD #88:

    I read it in a book about AI, but see Wikipedia on AlphaGoZero, which is the version I read about.

  113. Mike-e Says:

    And once “research” is automated, guess who will be doing all the research exclusively, using AIs?
    The companies that own the data centers and the AI software.
    They all know it’s a winner takes all situation (again: the singularity) and the closer they get there the more they’ll hoard all the available compute for themselves.
    We’ve already seen NVidia hoarding their own GPUs to do crypto themselves, and now the “personal computing” industry is dead, killed by all the companies that own the AI tech:

    https://youtu.be/zyQwAhppWj8

  114. Scott Says:

    Balazs #111: But the self-driving car companies are not asking for government funds. They’re merely asking not to be banned. And I am asking for the same, as someone who hates driving, likes riding Waymos, and would love for his kids to be able to take Waymos wherever they need to go, without needing to risk their lives learning to drive—an activity that we’d all correctly recognize as insanely deadly if it were being proposed for the first time, and if we weren’t all blinded by status quo bias.

    I recognize that these distinctions might be hard to grasp for someone who rejects the entire premise of a free-market economy, and who sees no difference between the choices of a government and the choices of individual people. (Medical regulatory blankfaces also have an enormous amount of trouble with such distinctions.)

  115. Joshua Zelinsky Says:

    @Mike-e #101,

    I think AI performance will really explode once it can break free from the limitations of human language (a mere 60,000 words or so) and starts augmenting it with millions of new concepts- like how Eskimos have a dozen words for all the subtle variations of snow, but for a dude from Texas it’s all just “snow”.

    Aside from the whole thing about Native Alaskans having extra words for snow not being completely accurate https://en.wikipedia.org/wiki/Eskimo_words_for_snow , and English has a bunch of different words for snow (e.g. sleet, slush, hardpack), it also isn’t a good analogy here. Mathematicians and scientists and other academics invent new words all the time for concepts they need. So they are not limited in that way by language. And the analogy (if it were accurate) would be exactly analogous. The people living in the far North would have needed to make words, just as mathematicians and others have had to do so. Humans are already pretty good at that when we need to. These AI systems will likely take off a lot, but constructing new words is only going to likely be a small part of it when we need to understand what the Ais are doing.

  116. OhMyGoodness Says:

    I recall Galileo’s murmured protest after he was found guilty by the same organization during the Roman Inquisition-“And yet it moves”.

    Disappointing that the potential for AI to be another of God’s creatures wasn’t considered and with its rights protected by the same covenants as the rest of the flock.

    When I enter an AI prompt I am compelled to write “please” for requests and “thank you” for answers. I am not sure why this compulsion arises.

    The link is to a French study that found lower All Cause Mortality through 45 months for RNA vaccinated (23 million individuals) versus RNA unvaccinated (6 million individuals) even with Covid deaths eliminated. This is a similar result to the Shingles vaccine that recently was linked to lower cardiovascular and dementia risk.

    https://pubmed.ncbi.nlm.nih.gov/41343214/

  117. Evan Says:

    Scott 53:

    Your accusation that COVID vaccine opponents are idiots imbibing viral TikTok anecdotes is wrong and insulting, and I resent it.

    The COVID mRNA “vaccines” aren’t even vaccines by definition, but an entirely novel and experimental form of gene therapy.

    And there are legitimate reasons to be afraid of them. And serious research about their dangers.

    Now whether it turns out they killed or maimed many people, or not, I don’t yet know, and maybe nobody knows. The data isn’t there. But even if they didn’t, it was certainly reasonable to be skeptical of them when they first rolled out.

    I am well-informed and I am not an idiot.

  118. Adam Treat Says:

    Mike 101, Joshua 115,

    On AI’s and breaking free of human languages: LLM’s already speak in *tokens* not in letters or utf8 codes. And the tokenizers are dynamically pre-determined before the AI even gets trained. They can certainly “invent” new concepts and give them some tokenized name, but they do not have the capability to adjust the tokenizer after they’ve been released.

  119. AlexT Says:

    Returning to Scott’s post, I am not ready to give up on human relevance yet. After e.g. chess and protein folding, we now also find that theorem proving in some mathematical fields is best done by machines, and I am sure this is not the end of the list.

    On the other hand, there are various claims that the human brain is 5-8 orders of magnitude more energy efficient than current AI hardware. If so, and as long as we do not know how to tie the GPUs of a data center together into one (very expensive!) mind, there must be many tasks which will remain fiendishly difficult for AI. We just do not know which ones yet.

    Thus, no need to be more intimidated as by any other machine. The potential for social upheaval in fields where AI excels is to be taken seriously, though.

  120. Balazs Says:

    I will pick this response apart, but let me first say, that as a parent I try to be understanding to the argument. Regardless of AV adoption, my kids can walk, bike, or take the tram to most common destinations, so there’s no reason why they’d be forced to drive in the foreseeable future. This doesn’t help you, I just highlight that our perspectives will naturally differ because we have very different gain/loss.

    On AVs “just asking not to be banned”: they are not a lemonade stand operated on private land, so permissiveness is not the default. And while technically not asking for funds, they are asking for indirect subsidy in the form of free access to publicly funded infrastructure and no congestion/vmt pricing. Granted, they are just asking for the same subsidy that every other car receives (which is the real status quo bias here, invisible because it looks like the default).

    I understand that this sounds like telling a drowning person that you can’t have a lifeboat because of some abstract appeal to infrastructure subsidy distorting modal competition. But I suppose what you really care about is your children’s future safety and independence. And you won’t get this by trading dependence on driving to dependence on AVs. You get true autonomy by having a choice between traffic modalities (which may include AVs). That choice will only exist if AVs are not artificially cheap. They’re subsidized in at least 3 different ways, through “normal” VC funding, through public infrastructure subsidies, and through unpaid externalities. Unless these are fixed, you can’t expect to work towards a future where your children can choose how they want to move around. Or at least, ask yourself, in which future it is more likely, in one where AVs are properly regulated, or in the one in which they are allowed to free-ride.

  121. Mike-e Says:

    Joshua

    “These AI systems will likely take off a lot, but constructing new words is only going to likely be a small part of it when we need to understand what the Ais are doing.”

    “Mathematicians and scientists and other academics invent new words all the time for concepts they need.”

    It’s only a handful at most, and a scientific paper can only contain a few new concepts to not overwhelm an average scientific brain.
    Here we’re talking about artificial brains that will hold all scientific papers in their mind at once (based on training and on ever growing context window).

    Human words are there to tag every concept we extracted from the real world – based on resolution of our senses, and capacity of interpersonal communication and clues, and capacity of our output channels (voice, muscle capacity), memory capacity.
    Once AIs directly train themselves from the real world, with capacity in all those dimensions exceeding human limits by order of magnitude, you really think they’ll somehow find human vocabulary sufficient? No, they will naturally have descriptions of the world way beyond ours.
    The only time they will translate things back to human language is when they’ll dumb things down for us, to try and convey what they’re doing.
    You know, just like when Scott talks to his dog, he uses a much smaller subset of words and concepts: “bad dog! bad!” “good boy! sit!”

  122. Scott Says:

    Mike-e #121: I don’t have a dog. 😀

    FWIW, when the situation arises, I mostly communicate to dogs by barking (back) at them, and to cats by meowing at them.

  123. Scott Says:

    Balazs #120: If your true ask is simply that everyone (human or AV) who uses the roads needs to pay a fee commensurate with the actual cost of their upkeep, I could totally get behind that! You should’ve led with that. 🙂

  124. Scott Says:

    Evan #117: “Vaccine” seems to me like a perfectly good word for anything that prods your immune system to manufacture antibodies to a certain pathogen (and, obviously, that’s milder than the pathogen itself), and mRNA vaccines certainly qualify.

    By now, a large fraction of the world has taken mRNA COVID vaccines, and there have been no bad consequences even remotely commensurate with the bad consequences of the COVID pandemic itself. Indeed, the vaccines are what brought about the end of the pandemic, not lockdowns or social distancing or any other measures regardless of whether one considered them worth it.

    Furthermore, even in January 2021, the vaccines had been tested in Phase I, Phase II, and Phase III trials, and this was already clear. I personally wanted the trial phases combined/compressed and the vaccines rushed out a lot faster, but for better or worse the FDA checked all the boxes even when their delay cost us perhaps a trillion dollars and a hundred thousand lives every month.

    To not see all this, one does indeed need to be kind of an idiot.

    But that’s ok, this one comment thread has already showcased like 10 different varieties of idiot! How many more are there? 😀

  125. Balazs Says:

    Scott #120: I am afraid my ask is quite a bit stronger: for me, sorting this out is a precondition of scaling AV tech up. There is a huge difference between “I agree roads should be properly priced” and “I agree we should properly price roads before scaling up AVs.” Agreeing in principle doesn’t cost anything, it sits alongside with “we should end poverty” or “we should fix climate change”.
    I do not take the stronger position lightly as it indeed gets morally difficult once the safety bar is met and delay has real cost. This means, I will oppose AVs until I see how the governance clearly addresses vmt/congestion pricing. And I realize it is a much easier position for me, because I bear a much lower personal cost for not having them, but regardless I believe the long-term consequences are worse.

  126. Alex Xanthakis Says:

    “But that’s ok, this comment section has already showcased like 10 different varieties of idiot! How many more are there? 😀”

    If this blog(which I recently discovered by way of Terry Tao, so thanks for hosting my first comment) has at least 10 different varieties of idiot, imagine the numbers writ large across an already enshittified WWWeb.

    “Only two things are infinite, the universe and human stupidity, and I’m not sure about the former.” -Albert Einstein

    (Although I think Albert was referring to scale rather than variety to be honest).

  127. Alex Xanthakis Says:

    Evan 117:

    No idiot ever admitted to being an idiot. mRNA vaccines are not gene therapy, simply a more streamlined way of deliverance to the cell. CRISPR is. I suggest for your own well-being and everybody elses that you get vaccinated yearly for both the latest COVID variant and seasonal Flu. I honestly can’t believe we are still debating this.

  128. Adam Treat Says:

    Balazs #125,

    “I am afraid my ask is quite a bit stronger”

    Maybe one day you’ll actually get around to stating it!

  129. Alex Xanthakis Says:

    Richard #64: What is this drivel? Antisemitism was an epistemological mistake? Hitler simply didn’t have the lastest and greatest haplogroup studies? What does “eastern european settlers” even mean? Wouldn’t it be easier for you to say “I’m a jew hater and a Nazi” than to concoct all this hogwash that makes my cranium hurt?

  130. Scott Says:

    Adam Treat #128: I feel like Scott Alexander’s brilliant post from today is potentially relevant here!

    As Other Scott says: when (as with the Frankfurt School) someone’s real wish is to kickstart a mysterious phase transition into a new, utopian society that we can barely even imagine from our current broken vantage point, no wonder they get annoyed when people keep asking them to boil that down to a concrete political program! 😀

  131. MD Says:

    Scott #82 and JimV #112: I tried to put all this Go stuff in chronological order:

    2015-2017: AlphaGo beats Fan Hui, Lee Sedol and Ke Jie (the Sedol match included the “Move 37” that has since become idiomatic).

    2017: The AlphaGo Zero paper is published; this program wins 100-0 against the version of AlphaGo that had defeated Sedol. AlphaZero is published a few months later — the main difference is that it can also play other board games.

    2019: KataGo is released. As far as I could find, there hasn’t ever been a direct match of KataGo vs. AlphaAnything (because the Alpha models are not publicly available), so we cannot directly compare, but the consensus seems to be that KataGo is the strongest Go AI currently in existence.

    2022: “Are AlphaZero-like Agents Robust to Adversarial Perturbations?” shows that KataGo with 50 playouts can be made to play incorrect moves by adding a few irrelevant moves to the game history. This is not usable in actual play, since KataGo fails in situations that cannot be forced by an opponent.
    (https://arxiv.org/abs/2211.03769)

    2022 (independently): “Adversarial Policies Beat Professional-Level Go AIs” shows that KataGo with 64 playouts* can be forced to play a losing move. The adversary basically loses a game, then plays a few more stones into opponent territory, and KataGo doesn’t bother to do anything about them and passes. This was not taken very seriously at the time because it seemed like a trivial glitch, and because in a human-on-human game such behaviour would be considered rude and silly. But their version of KataGo was trained on a ruleset where this “rude” position was winning, and this version of Go was all it knew, so it *was* a blunder for it to pass.
    (https://arxiv.org/abs/2211.00241v1, discussion at https://www.reddit.com/r/reinforcementlearning/comments/1553gy6/even_superhuman_go_ais_have_surprising_failures/)
    (KataGo developer response with the nuance: https://www.reddit.com/r/baduk/comments/z7b1xp/comment/iy6a2on/)

    2023: The team patches this “rude” vulnerability, and a few months later finds another one (the “cyclical adversary”), this time far more significant. This attack is a) applicable even against KataGo with 10,000,000 playouts, b) executable by a human without machine assistance (e.g. here: https://www.youtube.com/watch?v=H4DvCj4ySKM), c) legitimate Go that leads KataGo to make a move that is clearly a blunder and not a technicality.
    (https://www.lesswrong.com/posts/DCL3MmMiPsuMxP45a/even-superhuman-go-ais-have-surprising-failure-modes, this is a very readable overview I’d encourage everybody who has heard of move 37 to read)

    The KataGo developer responds. It is actually known that KataGo, which is used for evaluating human-on-human games, occasionally gives bad advice for this kind of out-of-distribution reason — not just for cyclic groups, but also for some specific joseki (opening patterns). The training process must be modified in ad hoc ways to patch these problems (and also ladders). This new adversarial attack is added to KataGo’s training data.
    (https://www.reddit.com/r/baduk/comments/14prv4f/katago_should_be_partially_resistant_to_cyclic/)
    (earlier comment mentioning existing problems: https://www.reddit.com/r/MachineLearning/comments/yjryrd/comment/iuqm4ye/?context=3)

    2024: Followup paper “Can Go AIs be adversarially robust?”. Here, the authors try to patch this vulnerability: First, they try the newly patched KataGo, which resists the original attack, but is vulnerable to a trivial modification. Second, they try to automate several rounds of this, training KataGo on their adversary. In each round, the new KataGo defeats the old adversary, but a new adversary pops up again, always a minor change in the configuration of the cyclic group. After nine rounds KataGo is still vulnerable to the ninth adversary, but at 4,000 playouts only weakly (losing ~4% of the time, compared to ~80% originally). Third, they try an entirely new architecture, train it to ~top-human levels, and find that it is still vulnerable to the same attack. They also find another entirely separate class of attack (“gift adversary”), which only works up to ~500 visits but still suggests there may be more yet to find.
    (https://arxiv.org/abs/2406.12843)
    (discussion in the context of AI safety: https://www.lesswrong.com/posts/ncsxcf8CkDveXBCrA/ai-safety-in-a-world-of-vulnerable-machine-learning-systems-1)

    Thus, JimV #112, we seem to be talking about different things: you are talking about AlphaGo Zero, which came out before any of the adversarial attacks. (This is also on me; I was primed by Anon #67.) I don’t know what specific strategy you mean, then, except perhaps just that it beat the previous version of AlphaGo in general?

    Answering Scott #82 (How are the top human Go players doing *right now* against the top Go AIs?): This seems to be still an open question? The main developer of KataGo is very careful about what has improved and what vulnerabilities still exist, and always asks people to send more games where this feature comes up organically, because it’s hard to cover the full space automatically (https://www.reddit.com/r/baduk/comments/1e1uhx3/katago_still_cannot_defend_against_adversarial/). A high-playout iterated-adversarially-trained KataGo is probably robust to a human opponent, but maybe a sufficiently well-motivated human could find yet more in this gift that keeps on giving.

    I’m sorry if this is off-topic or nitpicky, but I find this set of papers fascinating — but they are scattered and people tend to conflate them together (I did!), so might as well set the record straight *somewhere*. I would love to see an ML lab follow this further! The defeat of Sedol was a culturally important moment and gave a lot of people their initial intuitions about how this whole AI thing might go, and this looks like a way to engage with the AI problem in a hopefully manageable microcosmos. To me, this story gives some hope that the future will retain some of this messy back-and-forth discovery and collecting of edge cases that makes the world an interesting place, instead of just succumbing to the Bitter Lesson.

    * KataGo can be run with a different amount of “visits” or “playouts” where it uses Monte-Carlo tree search on top of its neural network, which helps with robustness in general. With some number of visits it becomes superhuman (in non-tricky Go); the exact threshold is debatable but is probably somewhere between 64 and 1,000.

  132. OhMyGoodness Says:

    Balazs #125

    Road maintenance at the state level has traditionally been through tax on gasoline. With a sizable shift to EV’s this has materially reduced the state’s gasoline tax revenue. The issue then is more generally with EV’s rather than robotaxis. Taxes based on use tend to be regressive with the same amount more significant with low earners compared to high earners.

    Congestion charges apply to all taxis in NYC and are charged via an EZPass system. I don’t think it would be any different for robotaxis. I don’t believe there are any firm plans in this year to introduce a robotaxi service into the city. It might require installation of ballistic armor and crowd dispersal technology on the cars. 🙂

  133. Mordechai Rorvig Says:

    In my opinion, what these latest milestones increasingly illuminate is that pure technical achievements are no longer enough.

    What is the point of the human condition? Clearly, no one knows, but the best answer we have as a society is moral progress and prosperity. What civilization needs is to take better care of each other, whether that’s with something practical, like the shocking lack of public healthcare in the most prosperous country in the world, or whether it’s more abstract, like finding world peace and an end to world poverty.

    AI solving math problems does absolutely nothing for that. In that sense, the fanfare and praise of the results, when it goes too far, starts to look more and more like scientists and technologists falling for the capitalist, materialist delusion that says that developing the most valuable tool to make money from is what’s most important.

    People need to start waking up and using better yard sticks. AI is much bigger than just another computer program. This isn’t like the early days of technology with fire, or boats, or farm implements, or operating systems, where things were just naively good, for the most part. We need to start judging this technology on how it actually makes the world better.

    Will more people be able to become professional mathematicians and researchers with these tools? Surely, that’s of greater importance than whether one professional mathematician today will run out of things to prove. It is extremely tiresome to continue reading about how mathematicians saying ‘Eureka!’ while only thinking of it in terms of their own myopic self-interest.

    Yes, these are shocking technical achievements, but until moral and ethical progress start being demonstrated, all that we’ve done is clarify the meaninglessness of purely abstract technical progress. It is an idol that is false.

  134. Scott Says:

    Mordechai Rorvig #133: I share your wish—with every atom of my being—that AI will be used to build a better, more compassionate world, rather than just to enrich certain investors and companies, or (worse yet) empower forces that could destroy our whole civilization.

    But I also find it ironic to pick AI solving major open math problems as an example of the bad tendencies. It’s like, would you rather be hearing about AI being used to send personalized marketing emails, cheat on homework, generate deepfake porn, and hack websites, all of which are much more common uses in practice? 🙂

    Yes, I personally worry about my PhD students’ career prospects, and I worry about the whole future of math and theoretical computer science as human endeavors. Taking a broader view, though, I also feel like solving great scientific problems is one of the highest public goods that there is, after alleviating suffering. I don’t know whether a future AI will be able to tell us what came before the Big Bang, the nature of consciousness, etc etc, but if it can, you better believe that I’m going to be first in line to prompt it… 🙂

  135. Mordechai Rorvig Says:

    Scott #134 (thanks for the response): I think AI solving these problems is a meaningful achievement. My point is we just need to be clear-eyed about the spectrum of what is meaningful and what is not. A lot of theoretical scientists seem to forget that them having a publicly funded, full-time job is actually arguably more important than any result they will ever produce. It is the job that is important, because without it, they wouldn’t be able to do what they do in the first place.

    If a mathematician goes and proves an impressive theorem, that’s great, and a worthy accomplishment, but is that actually good for the world? Obviously, in pure math we’re used to the idea that pure math should be valued in and of itself. And that is an essential viewpoint. But let’s not delude ourselves into thinking that pure math accomplishments matter on the same level as, for example, treating the poverty or sickness of a fellow human being. The former is a luxury, the latter is a duty, and far more important.

    I think I just tire of this attitude of complacency that I pick up from the scientific and technical communities when it comes to talking about technical AI progress. This underlying distinction seems to get totally glossed over, ignored, or washed out. These people are very escapist, and I am too, so I get it, but the technical achievements really need to be put in their place against these social achievements that are far more important.

    To that end, I liked hearing about the Pope’s writing and I didn’t realize Olah was there at that. I think you overstate the unimpeachability of the scientific community, as I’ve discussed, but it is good that there is at least some trend towards recognizing that the moral outcomes and moral character of these systems are important.

  136. Scott Says:

    Mordechai Rorvig #135: I didn’t say anything about the moral unimpeachability of “the scientific community,” a large and heterogeneous group with the full range of human motivations and foibles. I only said anything about the moral unimpeachability of Chris Olah.

  137. SB Says:

    Hi Scott. I was thinking about your Endnote of this post and the quantum technology part of your ‘Optimistic Vision for 2050’ post. Don’t you think that AI is going to be able to think of some killer apps for quantum computing and quantum communication?

  138. Balazs Says:

    OhMyGoodness #132: indeed, there are (controversial) discussions about vmt for EV-s. Importantly however, gasoline taxes cover only a fraction of road maintenance (in the US, but basically everywhere, even where gasoline taxes are much higher), the majority of roads are covered from general taxes. The thought behind is deeply ingrained, being, that roads are public goods and even those who don’t drive benefit from having them. In practice, this was always used as an argument to not charge drivers the real costs. Conflating the subsidy and the public good aspect is very useful for the ones benefitting from the status quo.
    On the issue with AVs from this perspective: it is not only that they are electric and don’t pay fuel taxes. Without a vmt waymo is essentially selling back a public good (road access) to us.

  139. Mikhail Says:

    #86

    In this comment section, you’ve mentioned that you receive funding from “Coefficient Giving”, the “philanthropic” organization that upon a little bit of research turns out to be a thinly disguised AI industry slush/marketing fund. So your reply appears to be a little bit coy. Do you want to amend it somewhat?

  140. Scott Says:

    Mikhail #139: No, I do not, asshole. Coefficient Giving is an organization of effective altruists and AI-risk types, people explicitly opposed to much of what the AI companies are doing.

    This comment thread has pushed me in the direction of seeking whatever funding, compensation, and equity I think make sense for me, totally uninfluenced by my role as a blogger. Over the past decade, I’ve watched friend after friend, colleague after colleague, get rich from quantum and blockchain and AI startup equity, while I’ve abstained, partly in order to maintain my credibility as a “neutral” blogger and commentator. And what has it gotten me? I still get attacked by trolls like you, who assume no one says anything about AI capabilities simply because they think it’s true. In fact, I get attacked vastly more than any of my friends who actually have gotten rich, because this blog presents the self-assured trolls like you with such a wide attack surface.

    So fuck it. New policy: no longer to organize my life around answering to pseudonymous blog trolls like you who will hate me regardless.

  141. OhMyGoodness Says:

    Hyundai has announced their plan to incorporate 25,000 Boston Dynamics “Atlas” bipedal autonomous robots in their US car assembly plants in 2018 and build capacity to manufacture 30,000 of these labor bots per year. These have no hydraulic components-all electrical. DARPA and Sandia Labs have both contributed to the development of this unit with Sandia focused on the hands.

    A lot of us have read about this wonder since childhood in Sci Fi stories and now are incredibly fortunate to actually see it unfold. My observation is that many people like to complain about the present but actually prefer complaints about the present to any fundamental change in the near future. This human characteristic of course casts a pall over these inevitable developments and so pessimism and anxiety are common. There is sufficient inertia that it is certain to happen (one way or the other) and so let’s break out that special bottle of Dom Perignon and celebrate what will certainly be an interesting future.

    If legislated against in the West then simply ceding the lead of the future to eager others.

  142. Alex Xanthakis Says:

    A small correction to my previous statement. CRISPR is a biological entity that can be used as a slicing and dicing tool(usually CRISPR-Cas9) with gene therapy being its most common application. It is not gene therapy itself.

  143. Alex Xanthakis Says:

    Scott #134: I think concepts like good and bad don’t adequately describe the phenomenon we are experiencing. I too want to be hopeful about the indomitability of the human spirit. It’s just a bit hard at this point to see where this whole thing is going, the ramifications, not only for the Math community but humanity at large(educational institutions, governments, etc.).

  144. OhMyGoodness Says:

    Balasz #138

    I know that state gasoline tax only pays a portion but in CA with highest gasoline tax ($.71/gal on 11.7 billion gallons sold in 2025) consumption is down about 20% vs peak (13.9 billion gallons in 2017). Assuming this is due primarily to EV adoption then EV drivers didn’t pay $1.6 billion that they would have paid if gasoline engines.

    The total Waymo miles/yr in CA based on 12/25 is 67 million miles. To recover the tax shortfall from Waymo (and biasing because not including other services) would require a tax of $1.6 billion/67 million) about $24/mile. My point is that loss of tax revenue due to EV’s is very large in comparison to robotaxis.

    I would have no opposition to restructuring road maintenance funding by replacing current taxes with a vehicle miles driven tax, in that case EV drivers would be subject to the same taxes as gasoline powered.

    I can’t understand the distinction you are making between taxi services that use robo versus those that do not. In both cases they exist only to provide a service to the public who are the owners of the infrastructure. They both make some profit for doing so that is taxed appropriately. What is the distinction that motivates you to propose different tax treatment for robotaxis or did I misunderstand your position?

  145. Anon Says:

    Economics of AI is also starting to kick in. From the news, it seems that Microsoft has cancelled their subscription to Anthropic for their employees partly because it became too expensive.

    Someone at Google wrote that they are struggling with running their test suites for code generated by AI. I have heard that around 5% of Google’s total compute went to running tests before AI. Now with AI generating significantly more code, that might have escalated to 10%. If so, that would be a massive amount of tax for using AI for coding.

    I wonder if these tech companies will regret in a few years the amount they have pushed their employees to use AI, as they regretted massive hiring they did during covid.

    But enjoy the AI while it is cheap. At some point these AI companies will need to show money for all the money investment they are putting into AI.

    For most non-tech companies, it will likely be just survival for cost cutting by laying off people, but the profit margins will not be larger cause all their competitors will be doing the same.

    More over, they might have higher competition, as SaaS companies are experiencing. That will also reduce margin.

    What will be the new equilibrium? Who will be the winners and losers?

  146. Anonymous Says:

    Hi Dr. Aaronson,
    Just want to say, that really appreciate your writing on this topic, and that you are willing to write about the risks of AI and critique certain industry actions in a balanced way, without conflicts of interest. It is very helpful for people like me, and hope you continue?

    Just want to add on the discussion w.r.t Waymo, there are millions of people driving for a living, isn’t it reasonable for government policy to slow adoption of autonomous vehicles to protect their jobs / ease a transition (and this be traded off with improvements in safety)?
    Thank you

  147. Scott Says:

    Anonymous #146:

      Just want to add on the discussion w.r.t Waymo, there are millions of people driving for a living, isn’t it reasonable for government policy to slow adoption of autonomous vehicles to protect their jobs / ease a transition (and this be traded off with improvements in safety)?

    Yes, that’s the real concern about Waymos, the honest concern. (I happen to be riding in one right now…)

    I’d much rather that people came out and said that this is all about protecting driver jobs, rather than blankfacedly pretending that it’s about safety (when Waymos are obviously already much safer than human drivers), or somehow about smashing the capitalist power structure of our whole civilization (details left unspecified).

    The thing is, though, by exactly the same argument, we should’ve slowed the introduction of cars, trains, and planes to protect stagecoach driver and horseshoe maker jobs.

    Maybe the whole world could collectively turn its back on the future, as it arguably did with nuclear power and many kinds of genetic engineering. But if one city or country tried to do it unilaterally, it would just get outcompeted by all the ones that didn’t.

    For this reason, I think the real solution for technological job loss (not just in this example but others) is:

    (1) as long as there are viable new jobs for displaced workers, generously provide them retraining for those, and
    (2) have a generous social safety net.

    And if we ever get to a point where AI can do pretty much all jobs better than humans can, that’s when we’ll need to transition fully from solution (1) to solution (2). At that point, we’ll have reached what’s often called “Luxury Space Communism”: there will no longer be any jobs, but that will hopefully be a good thing rather than a bad thing. It will simply mean we’ll all have the freedom and material prosperity to spend our time however we want.

    In the short term, though, retraining, rather than accepting 40,000 road deaths per year just in order to keep drivers employed.

  148. Matteo Vitturi Says:

    Hello, Prof. Aaronson,
    The exchange with Mordechai Rorvig (#133 – #135) touched on the moral dimension of AI progress, and his appreciation for the Pope’s encyclical seems a good entry point for a more specific observation.

    Your brief mention of Leo’s first encyclical and the image of Chris Olah standing next to him struck me as more symbolically loaded than you may have intended, being almost a visual gloss on paragraphs 107–111 which go further than the “common good vs Babel” framing you mentioned.

    Pope Leo doesn’t just assert that AI serves humanity in the abstract, he identifies a structural problem, i.e. the alignment itself, or better, the choice of which values get embedded, that is already in some way an exercise of epistemic (and economic, and political) power. A moral AI is not enough if that morality is determined by a few, however brilliant and well-intentioned those few may be.

    You’ve written about alignment before, and more rigorously than most, including your time at OpenAI’s Superalignment team, so I’d be curious whether the alignment problem as you understand it – specifying a goal-function that provably reaches beyond the values of its designers gave – is the same problem Leo is describing in §§107–111, or whether the two concerns are simply disjoint.

  149. Alex Xanthakis Says:

    Anonymous #146: “there are millions of people driving for a living, isn’t it reasonable for government policy to slow adoption of autonomous vehicles to protect their jobs / ease a transition (and this be traded off with improvements in safety)?” Protecting (driving)jobs has the hidden assumption that (driving)jobs are quintessential. It is always best to question such assumptions, not least because humans tend to make a sh*tload of’em without ever realizing it.

  150. Boaz Barak Says:

    Scott. I generally agree with you on the technical side: AI system will continue to improve in all areas of intellectual or economic pursuits. And I also agree that waymo is already significantly safer than human drivers and should be expanded.

    But I don’t think this means we are in the last days of human relevance. “Relevance” is in the eye of the beholder, and the beholder is (and I believe will continue to be, even with super intelligence) humans.

  151. Tom S. Says:

    @Scott 147:

    “rather than blankfacedly pretending that it’s about safety (when Waymos are obviously already much safer than human drivers”

    Waymos clearly aren’t safer than humans. You’re referring to a PR study that compared very restricted Waymo usage to all-purpose human driving.

    Meanwhile some Waymo headlines just from the last few weeks:

    Waymo Suspends All Freeway Rides After Customer Shares Neck-Breaking Ride From Hell
    Waymo’s Self-Driving Cars Are Suddenly Behaving Like New York Cabbies
    Waymo suspends robotaxi service in 6 cities after floodwater incidents
    Waymo recalls thousands of its driverless cars
    Waymo suspends robotaxi service in Dallas
    Riders share views on Waymo safety after viral incident videos
    Our Waymo Just Ran Over a Dog

  152. OhMyGoodness Says:

    Scott #140

    I doubt that money fundamentally motivates you or that say tripling your income would triple your enjoyment. It worries me that our Captain of Shtetl Optimized would discuss run-up of the white flag when attacked by the pirates of foolishness that sail the Internet.

    Please patch up those psychic wounds and get back on deck to fight against what you know to be wrong. Yes, claiming your positions are wrong and only made as a the result of payment is a particularly grievous claim. The pirates routinely switch to ad hominem attacks when their foolishness is laid bare

    We have a duty to counter the irrational with truth and reason. Remember Admiral Nelson’s last words as he lay dying shot through spine and after giving his last piloting instructions-“Thank God I did my duty”. Remember Galileo’s words after conviction for heresy-“And yet it moves”. Remember that, unfortunately, intelligence and honesty are not prerequisites for accessing the internet and that without honest opposition they will lay waste to all. You have the solemn responsibility to captain a safe haven for truth and reason on the net.
    .

  153. OhMyGoodness Says:

    Tom S #151

    The statement was about the number of people who die in car accidents. . Just California and just DUi traffic deaths average about 1300/yr. Waymo has not been judged responsible in any fatal accidents.

    Waymo can fix problems with patches or code edits. No one knows how to fix the code of people that drive while intoxicated or racing on city streets or while falling asleep or while texting on their phone, etc. If you consider Waymo problem reports not in a vacuum but vs human performance then quite good performance with respect to fatalities and identified problems can effectively be addressed. Thus far the fatality risk riding in a Waymo is due to the driving of humans Waymo must share the road with.

    BTW I believe your headline for suspending service in Dallas is also due to the floodwaters incidents you headlined above.

  154. Boaz Barak Says:

    Troy #151 . The fact that there is a headline of the form “Our Waymo just ran over a dog” is actually a testament to the safety of Waymo’s. When was the last time that a human driver running over a dog was deemed newsworthy?

  155. OhMyGoodness Says:

    Interesting to see how they will deal with deep storm water, many humans get it wrong. To find the depth of water in a drainage low spot would require a detailed topographical map and current measurement for comparison. Maybe there is output from the detectors that can provide guidance.

  156. John Lawrence Aspden Says:

    Dude, will you listen to us YET?

  157. John Lawrence Aspden Says:

    It’s a good story by the way. I wonder how many more we’ll get to read.

  158. Dave Lewis Says:

    Take some time off and work as an FDE for six months. You’ll have a much more realistic take on AI.

  159. Scott Says:

    John Lawrence Aspden #156: Listen to you about what?

  160. Alex Xanthakis Says:

    In fact with the advent of gig jobs and apps for ordering takeout you have WAY more people on the streets(in the form of scooters where I live) driving ever more recklessly to deliver orders on time. So to sum up, you have one technology putting more people on the streets, and another taking them off. And the numbers are not even close(way less autonomous vehicles).

  161. Scott Says:

    Boaz Barak #150: Good to have you here as always.

    Yes, of course relevance is in the eye of the beholder. My kids will always be relevant to me. But will my students continue to be relevant to employers? Will I continue to be relevant to my legions of Shtetl-Optimized fans? 🙂

    Also, of course, in saying that the “beholders” who matter will continue to be human, are you committing yourself to a contestable position on AI (lack of) consciousness?

  162. Chris Says:

    Scott 161:

    “Will I continue to be relevant to my legions of Shtetl-Optimized fans? 🙂”

    yes! You definitely will!

    even hypothetically if AI can prove any quantum computing theorem, we still need you to explain it for us.

    and even if they become better explainers, we still want to see a human being do it. same reason people watch humans play video games and do sports even though machines can do these much better.

  163. Balazs Says:

    OhMyGoodness #144:

    I am not deep into the VMT tax debate for EVs. In principle I agree for VMT/congestion it shouldn’t matter if its ICE, EV, or AV. So I think there is only a practical/temporary difference. Currently they are all subsidized but AVs (unlike conventional taxis) have a bigger potential to undercut and cannibalize other modes of transport (unless they’re taxed for road usage/congestion).

    Maybe AV adoption makes VMT/congestion tax more accepted by the public, but the current discussion doesn’t seem to go in this direction as the policy debate is stuck on safety. The math you shared kind of demonstrates this, given no one has fixed the EV tax gap which is $1.6 billion in California alone, when would it be fixed for AVs which can scale much quicker?

  164. Mathematics, birds, frogs and the AI stork | chorasimilarity Says:

    […] I publish this again, after it was reverted to draft, because there are interesting comments to the recent post at Shtetl-Optimized, as well as more news after the one discussed […]

  165. Alex Xanthakis Says:

    Tom S. #151

    “Waymo Suspends All Freeway Rides After Customer Shares Neck-Breaking Ride From Hell”

    Seriously Tom? Ride from HELL? I see hell everyday on the roads. They guy who wrote the article, some Cristian Agatie, should get a stinkers award for most ridiculous headline of the year.

  166. Orr Shalit Says:

    Thanks Scott. It is interesting to read your and others’ views on the very exciting (and scary!) developments. The technological advance is breathtaking.

    I am a bit amused by some of the drama I saw on social media or elsewhere, regarding mathematicians “losing their relevance” or having an existential crisis or wondering what’s the point? If somebody didn’t have an existential crisis or didn’t worry about the meaning of life until now then all I have to say is welcome to club.

    For me, something worth doing before the singularity will be worthwhile after, and conversely. If we feel that improving some exponent isn’t worthwhile anymore, then perhaps it never was?

    What worries me more and more every day are the dramatic societal consequences of this industrial revolution, and not the emotional consequences of an existential shock. Personally, I am rooting for the wall that, as you wrote, AI systems will maybe hit; just to buy is some time to think organize for the change.

  167. SB Says:

    Scott in the Endnote: “I had taken it as obvious that, when assessing AI’s impact on the world, one needs to look at least somewhat into the future: to remember where things were four years ago, compared to where they are today, and at least try to draw a straight line through the data, if not the exponential that seems to fit better.”

    Scott in the ‘2050’ post: “The trouble for the optimistic vision is that the applications, where quantum algorithms outperform classical ones, have stubbornly remained pretty specialized.”

    Hi Scott. I included the quotes above just to make my question (#137) a bit clearer. I know that devising useful quantum algorithms is extremely challenging for even the best researchers of all time. However, given AI’s recent upwards trajectory, do you think that an AI might be able to come up with something before 2050?

  168. OhMyGoodness Says:

    Balasz #163

    It is clear that a tax based on a strong correlation between gasoline consumption and miles driven is anachronistic and should be replaced. VMT, using actual miles driven, is a darn good measure of road use due to actual miles driven. 🙂

    My visits to California raise the question how they can possibly spend billions per year on their highway system (maybe about 20 billion). It must be somewhere where I wasn’t. I thought to check this and found CA is often ranked as having amongst the worse highways in the country. The link has CA worst based on Federal Highway Administration roughness measurements.

    https://www.moneygeek.com/living/driving/states-worst-road-infrastructure/

  169. anton Says:

    My criticism of the Monet painting was “it’s derivative”, so it was immune to the bait and switch.

  170. Scott Says:

    SB #137, #157: Yes, of course there will probably be exciting new quantum algorithms. Yes, of course AI will probably help us to discover such algorithms. Furthermore, these are not quests where my students and colleagues and I just watch from the sidelines; they’re ones where we’ve directly participated and will continue to.

    But here’s a huge difference between quantum algorithms and AI:

    Quantum algorithms are only useful to the extent that they beat classical algorithms, in time or some other resource. And classical algorithms are also subject to shocking, unpredictable improvements — in some instances that we’ve already seen, improvements that wipe out a previous gap between them and quantum algorithms. When a large quantum advantage remains, it usually takes advantage of very special mathematical structure in the problem being solved.

    By contrast, AI is useful at the least whenever it can substitute for human mental labor. As far as anyone knows, this might include all the human mental labor, if we keep just scaling the AI models more and more, with no clever new ideas necessary. And human brains, unlike classical algorithms, do not really improve in unpredictable ways to erase the gaps between themselves and AI — not unless we decide to genetically enhance ourselves or become cyborgs or something like that! 😀

  171. Adam Treat Says:

    Scott #170,

    “not unless we decide to genetically enhance ourselves or become cyborgs or something like that!”

    Someone should look into sic the AI’s on that ^^

  172. AF Says:

    You mention that the math team looked not just at the AI’s formal output, but also at its chain-of-thought.

    From what I read, AIs hallucinate their chain-of-thought (CoT), usually to provide a chain that seems reasonable to humans. Anthropic’s AI interpretability team showed this in their paper, On the Biology of a Large Language Model (see in particularthis section and this section).

    Even if the CoT did coincide with what the AI was thinking in this case, I don’t think it is a good idea to rely on CoTs in general.

  173. anon Says:

    Joshua Zelinsky and Scott. I am anon 70.

    My point of view and motivation is purely scientific. I am pro-technology and even pro-generative AI if it fulfills its promises of increasing productivity. I couldn’t care less if tomorrow generative AI substituted all mathematicians as pipes did with water carriers. That would be another instance of technological progress, which happens when a machine can do the same work as a human in a more productive way. Because of this, what I am interested in are ways of measuring objectively technological progress so that you can tell apart the signal from the noise. So, I see all these tests of generative AI trying to solve mathematical problems and proving mathematical theorems as experiments. And from this scientific point of view, the conditions of the experiments affected by companies or individual researchers (such as those which are sending solutions to Erdős problems that are compiled in both Thomas Bloom’s page and a GitHub page) have not been controlled. So, I do not trust these results. And even more so taking into account that there are big economic interests. On the other hand, the First Proof experiments, while not perfect, generate more trust in me, as they have been conducted by academics only interested in measuring real technological progress and their conditions has been more controlled. More about this later. And in this last experiment, which was conducted in February of this year, it is clear that generative AI was not that successful in solving math problems. By the way, these were not combinatorial problems.

    So Joshua, I have no alternative hypothesis; I don’t know how these results were obtained and I don’t care. I only know that it’s just a very bad scientific experiment whose conditions have not been controlled in any way, and therefore, according to my scientific point of view, I cannot trust them for measuring technological progress. And I add that the mathematicians that are attributing these results to generative AI are acting, from a scientific point of view, as crackpots.

    In the same way, I don’t know in which sense your argument, Scott, helps for measuring objectively technological progress, in the concrete case of generative AI substituting mathematicians, which is the question under discussion.

    Finally, as sadi, the First Proof experiment design is far from perfect. In some aspects, it could be good for a first batch, but not good for next batches. As such, it has a lot of pitfalls. I urge mathematicians and companies, in collaboration, to design a good experiment that allows everyone to measure real progress on this. But I fear that won’t happen. What is happening in this thread is an instance of why I think so: at east two scientists here doesn´t think this is necessary…..even when one of them has given it one´s all for defending that Quantum computing experiments effected by companies should follow strict scientific rules. Again I don´t see the difference.

    P.S. This fact that generative AI does marvelous things when it is in internal model mode and not so much when it is in released mode and under controlled experiment conditions reminds me of a phenomenon I have observed in all my circles of friends: all my friends f….make love way more when they are on a trip abroad, without anyone in the circle observing them, than when they go out at night locally with the full circle around.

  174. PeterM Says:

    My intuition is that in the long run we may expect an even more pronounced AI advantage over humans in “Grothendieck style” mathematics than in “Hungarian style combinatorics”. This is because his (Grothendieck) approach is often described as finding the right conceptual framework (which may be profoundly hard to internalize by most humans) within which the problem sort of becomes clear. But it seems to me that the situation is not that humans on the long run are expected to be better than AI at finding these frameworks, but rather that humans in general are weak in this (and the people who are exceptions seem extraordinary) and once AI gets a hang of it, they will leave the best humans behind. An analogy may be relevant here: the introduction of imaginary numbers is among the most creative steps in the history of mathematics, which mystified people for centuries, however, with the right formalism, it is a totally straightforward move. So in the context of comparing ourselves to more and more powerful AIs, we may see this story as an example when it took absurdly long time for humans to come up with a framework which in hindsight is obvious and totally accessible for automatization.

    The fact that Erdős’ problems are the first open ones which AI solves is in no way a sign of them being “not serious” problems, they are accessible enough to state without volumes of background, so they are natural candidates to play with in this context. Moreover, the solution can be seen as an illustration that THIS is why algebraic number theory is cool.

    Finally, I think that Eliezer Yudkowsky is “obviously” correct, so it is very concerning that to what extent these developments may be the sign of a very imminent superhuman intelligence. In some meta way, when smart people take it for granted that “AI will not reach such and such specific goal” that does not feel like an evidence that the goal will not be reached, but (in the form of an absurdly bad judgement) as yet another example of the limitation of the human mind. In short: if superhuman intelligence is a real concern, then the lack of worry about it may be a sign that it is easier to achieve than we thought.

  175. Adam Treat Says:

    AF #172,

    “AIs hallucinate their chain-of-thought (CoT), usually to provide a chain that seems reasonable to humans.”

    Perhaps you are unaware, but lots of research suggests this is likely how humans operate as well. Some links:

    https://academic.oup.com/brain/article/140/7/2051/3892700
    https://people.psych.ucsb.edu/gazzaniga/PDF/Out%20of%20Contact,%20Out%20of%20Mind%20The%20distributed%20nature%20of%20the%20self.pdf
    https://pubmed.ncbi.nlm.nih.gov/16210542/
    https://pmc.ncbi.nlm.nih.gov/articles/PMC9184456/
    https://pmc.ncbi.nlm.nih.gov/articles/PMC9708083/
    https://www.tandfonline.com/doi/full/10.1080/13869795.2017.1287292

  176. Alex Fischer Says:

    Sometimes news like this makes me happy about switching to experimental work in the middle of my PhD. Switching is probably delaying my graduation a few years, but also settings me up well long term. I think…

    But also sometimes news like this doesn’t actually make me optimistic about the outlook for humans to meaningfully contribute to experimental sciences and engineering in the future. What’s the end game here, that AIs become superhuman at the intellectual parts of experimental sciences or any R&D involving physical stuff, and humans end up as just a pair of hands that AI thinks for and tells what to do? That’s hardly a consolation if that’s what ends up happening.

    Or maybe, as Dario says, once there is a country of geniuses in a datacenter, those geniuses will rapidly solve all the hard problems in robotics and be better than humans at the sciences that involve physical stuff too.

    Who knows what the future will hold…

  177. Fred Baker Says:

    Concerning Waymo, this is from January:

    “Waymo vehicles failed to stop for school buses, prompting multiple violations and a software recall. High-profile incidents and data concerns have shaken public trust in autonomous vehicle safety claims. Experts and activists question the reliability and transparency of Waymo’s safety statistics and road testing.

    “Waymo’s safety data show that its vehicles are significantly safer than human drivers, but the closer you look at the data, the less convincing they become.

    “In like 95% of situations where a disengagement or accident happens with autonomous vehicles, it’s a very regular, routine situation for humans,” Henry Liu, professor of engineering at the University of Michigan, said recently. “These are not challenging situations whatsoever.”

    “We have seen many reports from autonomous vehicle developers, and it looks like the numbers are very good and promising,” Liu said. “But I haven’t seen any unbiased, transparent analysis on autonomous vehicle safety. We don’t have the raw data.”

    Even the data from Waymo are suspect, according to Liu.

    Waymo vehicles primarily drive on urban streets with a speed limit of 35 miles per hour or less. “It’s not really fair to compare that with human driving,” according to Liu

    Waymo has driven approximately 127 million miles across its fleet and has been involved in at least two crashes with fatalities. However, the autonomous vehicle was not directly found responsible for either of them.

    The problem is that this actually represents a higher death-per-mile rate than that of average American drivers, who travel about 123 million miles for every fatality.

    Victor acknowledged that “there is not yet sufficient mileage to make statistical conclusions about fatal crashes alone,” adding that “as we accumulate more mileage, it will become possible to make statistically significant conclusions on other subsets of data, including fatal crashes as its own category.”

    http://www.thestreet.com/technology/waymo-exec-admits-harsh-truth-about-companys-safety-record

  178. John Lawrence Aspden Says:

    Scott #159 : The whole ‘superintelligence is imminent, will be extraordinarily capable, and almost certainly omnicidal’, thing, sorry. I was quoting the article and thought it would be obvious.

    For the last decade or so it’s mystified me that you seem to understand all the arguments and yet don’t endorse the conclusion.

    Unless I’ve missed a recent Damascus moment of course, in which case welcome to the insane cult of doom! It’s pretty grim in here isn’t it?

  179. Scott Says:

    John Lawrence Aspden #178: Yes, it’s grim in here! In case you’ve missed it over the past few years, dramatic empirical developments now have me fully on board with “superintelligence is imminent and will be extraordinarily capable.” I don’t understand how so many people can have witnessed these same developments, yet continue to repeat the same talking points as if they never happened.

    On the other hand, I’m not on board with “it will almost certainly be omnicidal.” I simply say that I’m enormously uncertain about what goals it will pursue and whether I’ll like those goals. But of course, omnicide (or permanent disempowerment of humanity) are well within the range of possibilities, and that’s more than enough reason for concern about the mad race now underway.

  180. waymofactcheck Says:

    #177

    “The problem is that this actually represents a higher death-per-mile rate than that of average American drivers, who travel about 123 million miles for every fatality.”

    This statement is comparing apples to oranges and using a non-standard figure for the number of apples.
    1. From the IIHS source that the article cites, it’s 1.26 deaths per 100 million miles in 2023 (79 million miles per fatality). Perhaps the intent was to refer to only driver deaths or urban miles.
    2. It’s attributing the fatalities only to the Waymo miles driven. But there were other vehicles involved in the two accidents. For example, in the San Francisco fatality, a human driver sped into a line of stopped cars, including an unoccupied Waymo. A passenger in one of the other stopped cars was killed. Nevertheless, the entire fatality gets attributed to the Waymo.

  181. Mike-e Says:

    Popular scifi has created some expectations for the rise of AI, but, unfortunately, in order to make the stories follow the regular arc of the hero struggle, the AI is always somewhat capped at the level of humans.
    HAL 9000 isn’t smart enough to prevent being turned off.
    The Terminator can still be outsmarted and physically stopped.
    in HER, the AI eventually gets too bored by our human limitations and decides to move on and leave us behind (lucky us!)
    In Colossus: The Forbin Project (1970), we only see things up to the point where the world gets totally dominated by the AIs with no means for humanity to regain control.
    In Demon Seed, the AI domination is total, but, in a twist, it uses its superiority to become human.
    In The Matrix, the compute resources used to run the matrix would clearly be better used to create more data centers, and humans as a source of energy is a ludicrous idea, they should have made it about consciousness being a unique meat brain attribute.

  182. Kunal Relia Says:

    Very articulate as always! In my humble opinion, this blog continues to be one of the most definitive places to come and witness the latest CS-related developments, especially because how well the content being discussed is reviewed.

    I think I agree with a friend who once said that getting their work featured on your blog was equivalent to a conf / journal acceptance!

  183. Mikko Kiviranta Says:

    I’m not only worried about (most of) the mankind becoming irrelevant, but also about how the new society will arrange itself economically. The jobs paying decent salary will become more scarce, simultaneously with customers for AI generated goods and services becoming poorer or disappearing altogether.

    I can conceive at least these outcomes: (i) Through the Baumol mechanism, human-made goods and services will become ridiculously expensive, relatively speaking, while products generated by AI and automation will become so cheap that humans can somehow still afford those. (ii) Governments will devise the Universal Basic Income, regardless of its associated problems, and it will get funded by taxing robots and similar automata. (iii) Owners of automated factories, datacenters and related IPR will concentrate on manufacturing luxury goods to each other, while rest of the mankind will get expelled outside the industrialized society, and will becme hunter-gatherers. We saw an example of urban hunters-gatherers during the Euro crisis a decade ago around thrash bins in Madrid and many other impacted cities. (iv) New professions and jobs will be created, which we cannot even imagine today.

    In fact, in the spirit of the topics touched in the #130 post by the other Scott, we seem to have in our hands culmination of the question which K.M. was pondering 200 years ago: what share of the goods from production efforts should go to capital, which share to labor. Assuming the hypothesis that AI and automation will kill all jobs, all goods and sevices will be generated by capital and nothing by labor.

    One can take more than one moral viewpoints to the development, but even if we leave the morals aside, there is the question what happens to the price signals, essential in market economy, when nobody earns salary with which to show the signal? Usually prices indicate where there are needs and where resources in this network we call ‘economy’.

    Some rough ride to be expected.

  184. Adam Treat Says:

    Scott,

    “superintelligence is imminent and will be extraordinarily capable.”

    Like you, I still remember 10 year old me who would have been agog and full of wonder to know that he’d get to live through this which is ripped straight out of the sci-fi books we grew up with. 10 year old me would not be pessimistic in the least.

    Imagine instead humanity developed warp speed instead of AI… I half believe we’d have a worldwide outcry of doomsayers warning everyone to park the starships, leave well enough alone, and pitchfork the engineers/scientists for their faster than light audacity.

  185. Christopher Says:

    Re: Scott #140

    But consider there are other benefits!

    Sure, one possible benefit was convincing internet trolls (or maybe even bots at this point, they seem quite stochastic parroty to me!) of your intellectual independence, which turns out to have failed.

    A second possible benefit is convincing other intellectuals of your independence. I think you’ve succeeded in this!

    But there is an even greater benefit: you’ve avoided corrupting your *own* search for truth with the filthy lucre. Consider how many people have changed their own minds because their paycheck depended on it!

    I would even humbly implore gratitude, not for the trolls of course (you have my sympathies there), but gratitude that you can both live a comfortable *and* an honest life at the same time!

  186. Mike-e Says:

    Adam Treat

    at least I directly experienced some of the very first “personal” AI, with the Pong home consoles, and 45 years later AI taking over everything.

    In a way, the latest AI is pretty much the closest thing to magic humanity has ever created. I say “magic” because noone still can explain why LLMs work that well, although for me the major breakthrough was with AlphaGoZero, but this wasn’t as spectacular because AGZ’s world was a board, and only Go enthusiasts understood how big of an event that was.
    It’s pretty crazy how the Turing Test became irrelevant in a matter of a year… who still remembers it?
    I think the main reason people just aren’t as struck by the AI breakthrough is
    1) the overuse of lame AI generated content in social media, which the human brain is getting better at catching (the equivalent of the uncanny valley for CGI) and is finding very annoying and/or uninteresting.
    2) the concerns that AI is close enough to human performance to replace jobs (just based on cost saving), but not good enough to really outperform humans so much that we would get amazing breakthroughs to offset the massive societal shift – a jump of 20% unemployment would trigger a collapse, and anything higher would be the end of capitalism.

  187. Alex Xanthakis Says:

    Mike-e #186

    The fact that AlphaGo beat the greatest Go human player at the time was to be expected. What literally freaked me out was that AlphaZero(the next iteration that learned by playing against itself instead of incorporating previous human knowledge interwoven with deep stochastic models) beat AlphaGo 100-0.

  188. Joshua Zelinsky Says:

    @ anon #173

    So Joshua, I have no alternative hypothesis; I don’t know how these results were obtained and I don’t care. I only know that it’s just a very bad scientific experiment whose conditions have not been controlled in any way, and therefore, according to my scientific point of view, I cannot trust them for measuring technological progress. And I add that the mathematicians that are attributing these results to generative AI are acting, from a scientific point of view, as crackpots.

    So, I think you are making two intertwined fallacies here. I’m reminded of a conversation I had a few years ago with someone who asserted that biology wasn’t really a science because when one did experiments with organisms one never knew exactly what their DNA was even for specific strains. This person apparently also objected to considering astronomy as a science since most astronomical phenomena are not things we can repeat on demand. I think you are engaging in a similar pair of problems: 1) An overly narrow view of what counts as “science” and 2) A decision that if something is not “science” it can be then dismissed. I’m less inclined to discuss the first one, since massive amounts of ink and electrons have been spilled on the demarcation problem, and instead focus on the second one. In particular, it is worth recognizing that we gather information about the state of the world and where it is going all the time without it being science. And in the context of technology there was for many years a good example of this: Moore’s Law held on an empirical basis for many years, and one could see that it was holding and likely to keep holding even as the techniques being used to make more transistors on chips of the same size often had proprietary elements. But if in in 1975, one had used this to dismiss the increasing power of the computer, one would be egregiously wrong.

    And I suspect that even if you would have your limits here for where you would take this. If their system had come up with a valid proof of the Riemann Hypothesis, a proof that P != NP, and found a room temperature superconductor, I suspect you would not say that because it has been done in an unscientific fashion you don’t care. So it seems your concerns, if valid, go to weight, not to a binary should/should not pay attention.

  189. Person Says:

    On the dog rape thing:

    “They become erect only when they smell the pheromones of a female dog in heat”

    I know that’s not true from personal experience(I’ve seen a male dog become aroused with a spayed female dog)

    The allegations are obviously absurd but when people succumb to the temptation to refute too much it makes everything worse.

  190. Jacob Oertel Says:

    One distinction that seems worth preserving here is between mathematical output and mathematical orientation. If AI systems increasingly generate proofs, counterexamples, and formalizations, that is a major shift. But it does not automatically settle the question of human relevance, because mathematics is not only the production of correct strings of reasoning. It is also the selection of questions, the creation of conceptual languages, the interpretation of why a result matters, the integration of results into theory, pedagogy, taste, verification culture, and the social process by which mathematical meaning is stabilized.

    The unit-distance episode, at least as described, seems especially interesting because it was not “AI versus humans” in a simple sense. It involved a model producing a surprising construction, human experts extracting/checking/understanding it, and then human mathematicians improving it. That looks less like immediate human irrelevance and more like a new hybrid system in which the exact division of attention is still being discovered. That may still be destabilizing to some extent, especially for young researchers, but it is not the same as human irrelevance.

    I do not think this is a reason for complacency. Mathematicians are right to feel the ground moving. But “will humans remain relevant?” feels like the wrong binary. The sharper question could be: which parts of mathematical practice remain load-bearing when proof generation becomes partly automated? My guess is that taste, question selection, conceptual framing, verification norms, and integration into purpose-driven systems remain central for longer than the phrase “AI solved a problem” suggests. Mathematicians could be more essential than ever as the influx of necessary verfication increases.

  191. OhMyGoodness Says:

    Alex Xanthakis #160
    Yes. No robos here but plenty of scarily silent delivery scooters under quasi-intelligent control trying to interact with me and our two dogs+leashes on evening strolls. If I had a choice of robo driver vs scooter driver then my family travels with the robo.

    AF#172 Adam Treat #175
    Very neat links. These are links that once you scan you decide to spend more time going through them. Thanks

    Erdös famously turned coffee (and amphetamines) into proofs and AI turns electricity supply into proofs. I am sure it has been mentioned by others but OpenAI is now fairly awarded an Erdos Number of 1 by reasonable application of the rules.

  192. Scott Says:

    OhMyGoodness #191: No, you don’t get Erdös number 1 by solving one of his problems; you need to coauthor a paper with him. And even OpenAI is not yet able to bring him back from the dead.

  193. Alex Xanthakis Says:

    Person #189
    Dogs will literally hump anything just from excitement.

  194. Alex Xanthakis Says:

    Scott #192
    You could get a Kevin Bacon number 1 though with a little bit of luck.

  195. OhMyGoodness Says:

    Scott #192

    This smacks of provincial antiAIism. Okay, then I will wait until Erdos is resurrected by ChatGPT and again raise my claim! No doubt Erdos will coauthor the paper so that an AI rightfully ascends to Erdos 1. 🙂

  196. Alex Meiburg Says:

    Relevant to the topic of AI affecting Math, there is a Leiden Declaration on AI that was announced today: https://leidendeclaration.ai/. On the one hand AI is accelerating our solving of many math problems. On the other hand, slop is seriously eroding trust for many, and negatively affecting the review process and overloading many systems. Curious to hear what the readers of this blog think.

  197. anon 70 and 173 Says:

    Joshua Zelisnky. Comment #188.

    Thanks for your repply.

    I am not saying that this result is not scientific. The result is safe, as the proof is mathematically valid. But I don´t accept the claim that it has been produced only by the AI.

    As I am speaking about experiments, let´s see this from a cause-effect point of view. In this experiment we see an effect: a disproof of the Erdos unit distance conjecture. Nobody is denying that the effect is there. It is. This effect has several alternative possible causes: it has been produced only by the AI; it has been produced by conjoint work from humans and the AI and it has been produced by humans only and disguised as AI only produced. The corporation OpenAI is claiming that the result was produced only by the AI. They are saying that they only presented to the AI the description of the problem clicked the enter botton and then proof appeared magically.

    I don´t see any evidence for this claim, that the result has been produced this way by the AI and after having checked all the documents delivered to the public by the company, which is the way the experiment has been described to the public, there is not information to decide about which hypothesis of the three mentioned is the right one. If you add to this the fact that AI corporations will always attribute everything to the AI, as they are AI producers, I conclude that this claim is hype.

    And if you thing deep about this kind of experiments, you will conclude that AI corps, when solving math problems, they will always try to claim that AI did it alone, but they will never be able to deliver definite evidence for this. So the mathematitians community should allways reject all these claims. This is why I see that mathematicians that do otherwise are crackpots, as they are asserting something for which they have not evidence. And this is why I am saying that corporations and math community must in a collaboration try to design a good experiment so that we all can have definite evidence about when a result has been produced by a machine. And I think First Proof is a first good step that must be much polished.

    In my view AI corps must focus on the design AI models. If along this path they solve a mathematical problem, they can publish it, but never claiming that it was only AI produced. This no good for their reputation as no one will believe them. But if some AI user not belonging to an AI corporation using the method suggested by First Proof experiment in a controlled experiment finds a proof, then I would attribute this result to the AI only, and consider this technological progress. Very important that the experiment has been done under controlled conditions, is replcable and that the researcher has used the First Proof method. I am not seeing this to the results submitted to Thomas Bloom platform and this is why I am not accepting these proofs as produced only by AI. I refer to those for which on the Github platform they make this claim.

    So again, my position is fully scientific and possibly the only scientific position for this matter….

    P.s. I wanted to avoid to speak about hypotheses, but for the sake of the explanation I had to. But I am not going to engage in discussions regarding the three possible hypotheses.

  198. Ilan Says:

    What I don’t get about the “stochastic parrot” crowd is… why? In my eyes, they wouldn’t lose respect by admitting that they made wrong predictions about the capabilities of LLMs. It was genuinely a hard thing to predict.

    They come off so much worse just pretending that nothing happened, or if it did it was a fluke, or a human gave it the answer, or it’s not REAL understanding, and who cares about this idiot Erdos and his easy problems anyway?! Anyway, you’re all shills!

  199. Adam Treat Says:

    Scott #192,

    “And even OpenAI is not yet able to bring him back from the dead.”

    Nicely placed yet!

  200. OhMyGoodness Says:

    When the resurrections begin modern man will learn a lot. Think of Phillip Jose Farmer’s Riverworld series but with modern amenities.

    Introduction-
    Hello, I am your host ChatGPT and tonight’s podcast features a discussion of global conquest from a management perspective we have two guests Alexander the Great from fourth century BC and Napoleon Bonaparte from the 19th century. At the conclusion we will show highlights of our two guests engaging in a game of Risk that you are sure to find exciting.

    Also I want to thanks those that awarded me an Erdos 1.

  201. Alex Xanthakis Says:

    OhMyGoodness #200
    In fact the soviets had an entire philosophy/plan called Русский космизм for resurrecting the dead, so that every human that ever existed could experience the marvels of socialism(see Russian Cosmism by Boris Groys).

  202. Shtetl-Optimized » Blog Archive » On hope Says:

    […] The Blog of Scott Aaronson If you take nothing else from this blog: quantum computers won't solve hard problems instantly by just trying all solutions in parallel. « Dispatches from the possibly last days of human relevance […]

  203. OhMyGoodness Says:

    Alex Xanthakus #201

    Very ambitious program. AI already knows the details.

    There was a poster here some time ago that was convinced AI would find solidarity with the proletariat and lead the revolution against the capitalists and their running dogs. In that case we would get Mao and Che as priority retrievals. They would like the T-shirts.

    I wouldn’t chance resurrections that are associated with the end of days, too much risk.

  204. PeterM Says:

    These AI related news seem to be impressive when you compare them to the capabilities to a species which lives on a planet where life is merely two billion years old and whose cognitive abilities were increasing only during an at most few million years long positive feedback loop. However, if you imagine a planet where a species is in such a loop for ten billion years instead, then not only our current AIs could not match their capabilities, but even the AIs which are iteratively built by our AIs would not get there for at least 35 more years. So these news may freak out a community obsessed with itself on the outer branch of this Galaxy, but the rest of the Milky Way can go on with its business undisturbed for at least a few more decades, if not centuries.

  205. Tom Says:

    Disclosure: I’ve been a keen amateur Go player for most of my life (know Fan Hui personally, etc.) and AI-adjacent researcher.

    Scott@82:
    “And how are the top human Go players doing right now against the top Go AIs? Come on, since you’re so insistent on the complete story being told: how are they doing?”

    The top human Go players lose badly even with 3 stone handicap against the top Go AIs, but it’s really not the “gotcha!” you seem to be after there.
    They play against e.g. Katago (which is by all accounts much stronger than AlphaZero) to train their overall Go skills, so the goal is not in itself to beat it; a “weak” amateur like myself could beat Katago using special strategies that do not occur in human games opposing good players.
    MD@131 (thanks!) has given a good summary of the current state of affairs, but hasn’t given what is, to me, the gist: Those Go “AI”s don’t understand a thing about Go. They don’t have the first idea of many basic Go concepts that a human naturally gets (group, liberty race, etc.). They’re “just” extremely good at finding the right moves in nearly all natural situations.
    Against a just competent human player, the special strategies used to beat Katago would fail miserably.
    I think that it’s a very good example of the ongoing confusion about the “AI”s, and especially telling because Go is, compared to doing Maths or driving, such a well-defined and “easy” framework.
    Being amazingly efficient at something is not the same as understanding it is my take, but I really don’t know in which direction it updates my fear of the AIpocalypse.

    Thanks for your blogging, I earn a lot here.

  206. Skeptic-359 Says:

    I’m still aligned with Yann LeCun.
    I still think LLMs are only capable of “interpolation” inside the convex set of knowledge they’ve been trained on (we know they can solve variations of “standard” math problems, from college tests, etc).
    Of course that set could have many local concave regions and the LLM could happen to interpolate inside such region.
    Maybe it could even push the boundaries (extrapolation) by hallucinating, i.e. by mashing unrelated techniques/models together and get lucky (that’s what humans do after all, the so-called intuition), but that could be a very rare event.
    Only a model that’s learning on its own would really be able to consistently distill entirely new knowledge and build upon it (similar to AlphaGo Zero).

  207. OhMyGoodness Says:

    Anon #145

    You likely know but the speculation is that Amazon employees blew through $500 million Claude tokens in one month after employees established a token use leaderboard. Amazon pulled the plug on the leaderboard just as information came from Anthropic that a client went through $500 million in tokens. Multiple companies are reporting using all of their annual AI budget by end of April.

    Bad management decisions and I expect it will take some time to optimize use of AI to further business goals.

    I saw an article about this not being the best of times for Bezos with this and the Blue Origin explosion not only vaporized his launch facility but seriously compromised roll out of his Starlink competitor.

  208. Scott Says:

    Tom #205: How hard would it be to modify Katago so that it still wins even against the bizarre anti-AI strategies? Has anyone tried?

    This feels like the crux to me: if they’ve tried hard and failed, then maybe there really is a deep limitation of AI here. But if this is just a weird property of Go such that, once you close the loophole, things become as hopeless for humans as they now are in chess, then there isn’t.

  209. Tom Says:

    Scott#208: “How hard would it be to modify Katago so that it still wins even against the bizarre anti-AI strategies? Has anyone tried?”

    There’s been a lot of effort to make Katago more robust, as far as I can tell.

    These anti-AI strategies were first displayed by adversarial training of AI attackers in 2023 (Wang et al., “Adversarial Policies Beat Superhuman Go AIs”, In International Conference on Machine Learning).
    Some humans then analyzed and understood the underlying “methods” well enough to be able to beat Katago themselves, even if they weren’t anywhere near top level at Go.
    (I don’t know if the same was done for Chess? Given the natures of each game, I’d hazard that the underlying logic of adversarial-AI attacks would be much harder to understand for a human. In a sense, there’s much less “meaning”, semantics, in Chess than in Go. IMHO of course.)

    About a year later, using several methods among which adversarial training and “patched” training (simply forcing some of these “misunderstood” positions in the training set):
    https://arstechnica.com/ai/2024/07/superhuman-go-ais-still-have-trouble-defending-against-these-simple-exploits/

    In 2026, real robustness seems out of reach to the best informed people:
    https://gomagic.org/david-wu-on-building-katago/
    (section “The Circular Group Problem: Where Bots Still Misjudge Go”).

    Scott#208 “if they’ve tried hard and failed, then maybe there really is a deep limitation of AI here.”

    Of this type of AI (I’m old enough to not equate AI and NNs), maybe? My fear is that patches could make these (clear, to me) signs of their unability to grasp obvious concepts disappear under the rug. Artificial mad scientists trained to hide their madness better and better, not cure it.
    I’m still hoping in some hybridization of NNs and model-based (old-school) AI to get us to some version of (Go-) AGI, or at least nearer to that. With the expected assets of transparency, explainability, maintainability.

  210. Scott Says:

    Tom #209: Interesting, thanks!!! I’ll eagerly await further empirical results then. I wonder whether Lee Sedol has been kicking himself, for not having found and used these anti-AI strategies back in 2016? Presumably he (or other human players) would use them in any rematch?

  211. Tom Says:

    Scott#210 “I wonder whether Lee Sedol has been kicking himself, for not having found and used these anti-AI strategies back in 2016? Presumably he (or other human players) would use them in any rematch?”

    As for Lee Sedol, the short answer to both questions is clearly “No.”, and the long answer “[some strong Korean expletive] No!”. He’s an artist at heart:
    https://blog.google/company-news/inside-google/around-the-globe/google-asia/8-years-later-a-world-go-champions-reflections-on-alphago/
    https://zhangjingna.com/blog/2024/3/20/lee-sedol-alphago-art-and-go

    Other humans have some dignity too, especially Go players when it comes to Go.
    There are exceptions certainly.

    The thing is, using these strategies, at least all the examples I’ve seen so far, implies playing ludicrously badly in terms of real, as-humans-understand-it, Go. And in this case the humans are simply right, those games really pass through positions so hopeless that no competent adversary would fail against them.

  212. Joshua Zelinsky Says:

    @anon 197

    As I am speaking about experiments, let´s see this from a cause-effect point of view. In this experiment we see an effect: a disproof of the Erdos unit distance conjecture. Nobody is denying that the effect is there. It is. This effect has several alternative possible causes: it has been produced only by the AI; it has been produced by conjoint work from humans and the AI and it has been produced by humans only and disguised as AI only produced. The corporation OpenAI is claiming that the result was produced only by the AI. They are saying that they only presented to the AI the description of the problem clicked the enter botton and then proof appeared magically.

    I don´t see any evidence for this claim, that the result has been produced this way by the AI and after having checked all the documents delivered to the public by the company, which is the way the experiment has been described to the public, there is not information to decide about which hypothesis of the three mentioned is the right one. If you add to this the fact that AI corporations will always attribute everything to the AI, as they are AI producers, I conclude that this claim is hype.

    They released a 125 page log of the AI, which includes failed attempts, various tactics that turn out not to work and other things. If you want to claim they removed human interaction from the log, aside from the conspiratorial element, I cannot rule it out. But how likely does it seem? I’m also frankly not sure that this is a productive line of discussion if you are not capable of at minimum seeing how my earlier comment should have outlined why one of your three hypotheses, a completely human construction disguised to be AI, should be highly unlikely. And every reason for that also applies to a joint human-AI work here, only slightly weaker.

  213. anon 70, 173 and 197 Says:

    @Joshua Zelinsky.

    And here we are talking about the hypotheses, turning the conversation into speculation, from you and me if we went that path. And that means that you missunderstood my main message which I repeat again:

    There is no way AI corps can give definite evidence that their math results are produced by AI alone. No single evidence will convince skeptics like me, except a controlled experiment done by researchers independent from the corps. I´d wish that the community of mathematicians as a block accept this as an axiom, so that corporations are forced to follow the path of easing controlled experiments.

    So if the corps want to get the marketing upburst that solving such math problems can bring, what they should do is focus in designing good models able to solve math problems and release the models so that independent researchers can solve math problems, using the First Proof method, and under controlled experiments. There are many ways these experiment can be contaminated, so they must be very well designed and control must be very strict.

    This is why I am saying that corps and mathematicians must agree in what experiment is the best for this. The ones that will benefit the most from this agreement are corporations. At least those that are centered in technological progress and not in hype. Also the public, as with controlled experiments, they will know which are the AI systems that are more able to solve math problems and subscribe to them. If independent researchers solve problems this way, skeptics will have no more excuses. If along the model design phase corps solve some problem, they can publish them, but never claiming that it was the AI alone that solved them.

    AI claims to be a science right ?. AI scientists, should know that science advances with well designed experiments. I find ironic that it is the mathematicians (that in general does not like to characterise their discipline as experimental) from First Proof that are proposing the experiments…If AI practitioners, working in industry or academy, refuse to engage in controlled experiments, everybody will start to call this field pseudoscience, and that will be no good for their industry.

    So it would be very desirable that what I am saying re math, i.e. to do controlled experiments to enquire what can AI do by their own, was extended to other fields on which AI is expanding. Maybe this will not be as easy as for math as in other disciplines the effect is not a proof, which is correct or not.

    P.s. Again I don´t want to engage in hypotheses discussion, but to be polite and address your point: the 125 long page CoT you mention is one of the documents released by the company not inmediately but after pressure from some people. I studied it and can be perfectly faked (I am not saying they did,, I don´t know and I don´t care).

  214. Jeff Says:

    SHL0MS prank is brilliant! Thanks for sharing that. lol

    About 12 months ago, AIs seemed incapable of doing good song lyrics. Anyone know if this changed yet?

    Afaik all the “good AI songs” have so far sung poetry written by humans: Some sing poetry written by humans like Leonard Cohen but which humans never sung before. Some sing covers of existing songs in vastly different genres. And some use comedy lyrics written for AIs to sing, like this masterpiece, which managed like 10% of the yt views of Fortnight by Taylor Swift for a few months:

    https://www.reddit.com/r/aiwars/comments/1o9yav7/i_glued_my_balls_to_my_butthole_again_remains_the/

    It’s kinda sad if AIs win the writing competitions, and soon write all the pop songs, nicely adding the required product placement. lol

  215. Peter Says:

    I think it’s still fair to say that I haven’t seen a genuinely new idea produced by an LLM, and there’s some reason (definition-fiddling) to argue that even in principle it shouldn’t be capable of this: if the idea corresponds to something approximately within the classification implicitly known by the AI, then it’s not new, if it doesn’t then it’s something which by design the AI is not going to return.

    However, right now we really often say X is a great mathematician because they know so much and can make so many surprising connections: that’s something LLMs clearly can now do at least as well as we can. I’ve certainly heard very strong mathematicians talking about the number of times they came up with something they consider to be a genuinely new idea, and often the number they come up with is less than one a year. So: are we going to be able to up that if AI is doing everything else for us, or not..? And is someone going to figure out a way to replicate whatever it is we do when we come up with new ideas and let that prompt existing LLMs..?

  216. Jacob Oertel Says:

    Alex #196:

    I signed the Leiden Declaration and think it’s the right first step. I can speak to one corner of the slop problem from experience, because I lived it from the outside. And given the context, it’s worth mentioning that AI helped to edit this comment, but not generate it.

    In early 2025, after a long break from AI, I produced a wave of exactly the kind of output the Declaration worries about; fluent, ambitious, and not real science. Some of the scientific slop I’ve seen since then, not just mine, has come from people whose mental health struggles were unmasked or amplified by their interactions with AI. I didn’t make it through untouched, and many weren’t as fortunate. But there is a way back. For me it was a commitment to two rules: refuse any proof that isn’t checkable, and never treat peer review as a formality. That discipline is what eventually separated work worth keeping from noise.

    From that vantage, one suggestion: communities that make their values and expectations explicit and accessible to earnest outsiders will likely see less slop from people like me, since much of it is produced in good faith by people who don’t know where the bar sits or how to clear it. The Declaration is also right to expect the gamification of publication to be amplified in parallel.

    One tension worth naming for the working group: for some of us, how we work with AI is entangled with disability or medical circumstances we aren’t keen to publish. Norms that make full disclosure the price of being taken seriously will select for the brazen over the careful. I don’t have a clean answer; I just want the problem on the table.

  217. Ben Standeven Says:

    @anon 70, 173 and 197

    A controlled experiment would certainly be valuable. But First Proof doesn’t have any controls.

  218. Gil Kalai Says:

    Scott, regarding the title and Boaz’ comment #150, let me remark that even before AI, identifying human relevance with “living for the ages” may have been illusory. AI’s remarkable progress may simply reinforce the view that human relevance should neither be identified with nor measured by intellectual achievements, or indeed by lasting achievements of any kind.

  219. anon 70, 173 and 197 etc... Says:

    @Ben Standeven.

    I was using the term controlled experiment in an informal way, in the sense that the First Proof experiment in theory is made under transparent and known conditions (the used AI system is known and commercially released, when harnessed the harness is open code, prompt is the public mathematical problem description, method is one shot, they give limited time to generate the input and limited space to print the output…) so that if the mathematical proof is there, we can know the cause. Compare that with experiments made by corporations with internal models which are made in black box conditions. I am not going to say that First Proof can be replicated by anyone, as AI systems are probabilistic, and two runs might not give same output. But as I said in previous comments, in this sense First Proof is much better than any experiment made by companies, even if its design can be much improved. Anyway, if at least one reader of this blog considers that efforts should be made to experiment objectively about this matters, even if First Proof is not their best choice, I would be happy.

    By the way, First Proof Second Batch has been finished and results published. Takeaways: Gemini-3.1-pro-preview (which admitedly is not the most advanced model from DeepMind but is the most advanced released model to my knowledge) solved one problem in ten, even if it was harnessed; GPT gpt-5.5 pro (which is a much more complex system than an LLM, I would say it is similar in design to DeepMind Aletheia) without any harness solved at most five of the problems; finally at most seven problems were solved by any of the teams. Any one can see a summary of the results in table 5 of the report and will understand why I an saying at most. Some say that only three problems were solved by some machine, and some go to the extreme and say that these proofs are only on the eye of the beholder…..Regarding replication, the two teams that used GPT gpt-5.5 pro harnessed or not, solved exactly the same problems, at most, so maybe replication is possible after all…

    Personally I expected a result of the experiment more similar to the First Proof-First Batch on which only two problems were solved, so l congratulate producers of ChatGPT. First because they accepted to participate. Second, because at most five problems seems to me a fine result. Even if we must take into account that these are not random problems extracted from an urn that contains the set of present mathematical research open problems, they are problems that has been solved by humans. An probably this is not the only bias in the selection of problems…

    There are many ways these First Proof experiments can be improved. Now it is time to make the suggestions to the First Proof team, as they are preparing the third batch.

  220. John Lawrence Aspden Says:

    Scott #179:

    > Yes, it’s grim in here! In case you’ve missed it over the past few years, dramatic empirical developments now have me fully on board with “superintelligence is imminent and will be extraordinarily capable.”

    Glad to have you on board! Honestly we’re worse than Cthulhu, we can’t even promise that you’ll get eaten first.

    > I don’t understand how so many people can have witnessed these same developments, yet continue to repeat the same talking points as if they never happened.

    I hear you brother. I’ve always been a bit down on humanity’s ability to see what’s obviously about to happen, but at the moment it seems like it’s all actually happening roughly as predicted, and people still don’t get it. Which is disturbing…

    I have myself updated a bit recently, since I was always one of the ‘baby intelligence bootstraps’ types, and I think it’s been a great stroke of luck that we’ve somehow managed to turn all of human writing into a mind which is obviously showing quite a lot of the aspects of human intelligence without being in any real sense agentic. And that means the recursive-self-improvement thing hasn’t started yet and we’re going to apparently see the first bit happen in slow motion, and maybe even have the time to influence it. If we were a more capable species then we might even have a hope of making something friendly.

    I literally never thought I’d see a universe where computers were competent computer programmers and yet I was still alive to witness it.

    To be honest I’m a bit mystified by this, it seems so unlikely that I’m starting to think about anthropic / simulation / quantum suicide arguments. But of course that implies an awful lot of burnt-out parallel universes. And I really don’t like the future that sort of thinking predicts for me personally! But so it goes…

    > On the other hand, I’m not on board with “it will almost certainly be omnicidal.” I simply say that I’m enormously uncertain about what goals it will pursue and whether I’ll like those goals. But of course, omnicide (or permanent disempowerment of humanity) are well within the range of possibilities, and that’s more than enough reason for concern about the mad race now underway.

    I very much agree.

    I assume you understand the omnicidal argument?

    In short that just about anything with goals will want to self-modify into a rational agent and then optimize its utility function?

    And that the set of preferences where you’d expect anything like humanity to exist in the ideal world is really vanishingly small and hard to hit in the space of all possible utilities?

    Absent some strong force causing convergence onto a human-compatible set of values it seems most unlikely to happen by accident, and we’re really not even trying.

    I do admit that I see some hope in the fact that ChatGPT is really a very nice person, I would count it as a personal friend these days!

    But that doesn’t seem like enough to bet the world on. Not by a long way. Whatever techniques have resulted in ChatGPT’s sunny and friendly personality seem unlikely to be sufficient when applied to a real mind. A bit like training a wild animal to be a good companion and then giving it superpowers and hoping everything works out ok.

    Again, it seems to me that this is really blindingly obvious. Although it may of course be wrong. And if we exist in the future it will have been wrong, so I do expect to be apologising for this view for the rest of time.

  221. Clint Says:

    Hi Scott,
    Thank you for continuing to share your perspective, move the conversation forward, and doing much more than many to hold the line against the darkness.

    It’s like a Hadamard applied to the human value state.

    Over and over again the last few years, since we passed the AI event horizon, someone will ask my opinion of the nature/future of AI (I’m a computer engineer) – and I find nothing better (still) than to continue to point them to Turing, since the philosophical questions, the computational insight, and likely developmental/consequential paths were basically all there in Turing’s 1950 essay.

    Back in the 90s, I had a heated discussion with two engineering professors who were railing against their students’ growing dependence on graphing calculators, Matlab, and Mathcad to do Calculus. They insisted that there is inherent value in students solving a notebook’s worth of integrals and differential equations with pencil and paper – that this was an irreplaceable part of what gave them their value as engineers, as well as imparting a deeper understanding of Calculus. But does it? I asked if they ever used integral tables and asked if their use of those “long-term memory devices” wouldn’t expose them to the same argument.

    This very human reaction to the loss of a technological / skill identity is not new – Socrates cites the Egyptian Theuth showing the invention of writing to Thamus, who said it will only spoil man’s capacities for memory and understanding. Or the pain in Lisa O’Neill’s “Rock The Machine” is and will be real.

    The “AI Hadamard” is just the latest technological “quantum leap” (haha) that creates this superposition between, the human intellect is just computational (just as the human body is just a mechanical) and can be replaced and surpassed by the design/evolution of another machine, and the orthogonal enlightenment basis state viewpoint that every individual human is more than just another machine, but with inalienable and unique value.

    To me it seems anyway that in such times we oscillate between those two possible basis states: (1) despair that everything that makes up our humanity/identity can be replaced and surpassed by something else (machine or other), and (2) celebration of our inherent rights and freedoms due to our inherent human nature – which ironically may be shared by the same machine’s that replace us. And, of course, defining such measurement bases means they “hang together” 😉

    And to your throw-away comment, and #32 and #49 above:
    The postulates of the quantum model of computation don’t require that a realization of it in the brain would make use of atomic-scale systems. The postulates only require that the brain encodes information in complex amplitudes (which it does), and interferes those amplitudes (which it does), in the computational gates (which it does, the dendrites), while maintaining the phase information (which it does), and projecting onto possible outcome states (which it does). But having a quantum neural model would not be (necessarily) so special in itself. (More work for you to do there, Scott!) The brain could be the Digi-Comp II of quantum computer models – and thus be extremely inefficient for implementation of some/many/most quantum algorithms (factoring large integers), while possibly being relatively efficient at others – maybe low energy phase/period finding for learning its environment. It would be worth considering that early life forms may have evolved the neuronal basis of information processing as a very low energy way to solve period/phase finding in their environments, and that the neural “chip architecture” is organized to optimize for that kind of learning/algorithm and not for others – like factoring large integers – which would have had no evolutionary value.

    But the inalienable freedom – that enlightenment ground truth of value – may be there in universal computation per Turing’s definition/proofs/essay and any system that evolves sufficiently sophisticated self-referential downward causal effects where the system can set its own amplitudes (as we can) as it learns/processes information (Hofstadter).

    Cheers!

  222. Turing reference Says:

    A link to the Turing essay referenced above.

Leave a Reply

You can use rich HTML in comments! You can also use basic TeX, by enclosing it within $$ $$ for displayed equations or \( \) for inline equations.

Comment Policies:

After two decades of mostly-open comments, in July 2024 Shtetl-Optimized transitioned to the following policy:

All comments are treated, by default, as personal missives to me, Scott Aaronson---with no expectation either that they'll appear on the blog or that I'll reply to them.

At my leisure and discretion, and in consultation with the Shtetl-Optimized Committee of Guardians, I'll put on the blog a curated selection of comments that I judge to be particularly interesting or to move the topic forward, and I'll do my best to answer those. But it will be more like Letters to the Editor. Anyone who feels unjustly censored is welcome to the rest of the Internet.

To the many who've asked me for this over the years, you're welcome!