Movie Review: “The AI Doc”
Yesterday Dana, the kids, and I went to the theater to watch The AI Doc: Or How I Became An Apocaloptimist, the well-reviewed new documentary about whether AGI will destroy the world. This was surely the weirdest family movie night we’ve ever done. Firstly, because I personally know probably half of the many people interviewed in the film, from Eliezer Yudkowsky to Ajeya Cotra to Liv Boeree to Daniel Kokotajlo to Ilya Sutskever to Jan Leike to Yoshua Bengio to Shane Legg to Sam Altman and Dario Amodei. But more importantly, because this is a documentary that repeatedly, explicitly, earnestly raises the question of whether children now alive will make it to adulthood, before unaligned AI kills them and everyone else. So pass the popcorn, kiddos!
(We did have popcorn. And if the kids were scared — well, I figured we can’t shield them forever from the great questions of the world they’re entering. But actually they didn’t seem especially scared.)
I thought that the filmmaker, Daniel Roher, did about as good a job as can be done, in fitting into a 100-minute film a question that honestly seems too gargantuan for any film — the question of the future of life on earth. He tries to hear out every faction: first the AI existential risk people, then the AI optimists and accelerationists like “Beff Jezos,” then the “stochastic parrot” / “current harms” people like Emily Bender and Timnit Gebru, and finally the AI company CEOs (Altman, Amodei, and Hassabis were the three who agreed to be interviewed), with Yuval Noah Harari showing up from time to time to insert deepities.
Roher plays the part of an anxious, curious, uninformed everyman, who finds each stance to be plausible enough while he’s listening to it, and who mostly just wants to know what kind of world his soon-to-be-born son (about whom we get regular updates) will grow up in.
I didn’t think all the interviewees were equally cogent or equally deserved a hearing. But if any viewers were actually new to AI discourse, rather than marinated in it like me, the film would serve for them as an excellent introduction to the parameters of current debate (for better or worse) and to some of the leading representatives of each camp.
If I had to summarize Roher’s conclusion, it would be something like: go ahead, enjoy your life, have children if you want, but understand that now is a time of world-historical promise and peril much like the early nuclear age, so pay attention, and demand of your elected leaders that they ensure that AGI is developed in a pro-human direction, because tech leaders (even the relatively well-intentioned ones) are trapped in a race to the bottom and can’t get out on their own. Honestly, I’d have a pretty hard time improving on that message.
The main thing that gave me pause about the film was not on the screen but in the theater, which was nearly empty. For the film to serve its purpose, a significant fraction of the world will need to see and discuss it, either in the theater or on streaming. So, y’know, it’s still playing.
For whatever it’s worth, here were my wife Dana’s comments: “The biggest flaw of this movie is that Daniel Roher never breaks out of his ‘clueless everyman’ character, even when he’s talking to the most important people in AI. He wastes an opportunity to ask them non-superficial questions, questions deeper than ‘so, uh, are we all gonna die or not?'”
And here were my 13-year-old daughter’s comments: “So many of the people they interviewed seemed like hippies, who don’t know what AI will do any more than I know!” Also, after Daniel Roher wishes Sam Altman mazel tov on his forthcoming baby: “Sam Altman is Jewish?!”
And here were my 9-year-old son’s comments: “I thought this would be a movie, where AI would try to take over and the humans would fight back! I had no idea it would just be people talking about it. The documentary kind of movie is so, so, so boring.”
Follow
Comment #1 March 30th, 2026 at 12:20 am
My impression from the movie was that even the pessimists (maybe not Yudkowski) were pessimistic not because they thought that AGI would take over the world and decide to kill all the humans, but because they thought that there would be vast economic disruption, first because AI will be able to do much work better and more cheaply than humans, and second because wealth will become even more concentrated in the hands of a very few corporate entities. There was also a space-race concern that states we regard as rogue would leap ahead in building AI. But it all seemed to be about the fact that humans were likely to cause problems by misusing AI or being insufficiently careful in anticipating the effects of its deployment, not that AI was going to go off on its own to do bad things. As an accelerationist, my inclination is to say full speed ahead, and fix problems after they happen rather than hold back development for fear of what might happen, because then we’re going to miss out on the benefits; China builds vast networks of high-speed rail, and the US does not, because the US gates development through processes designed to retard it.
Comment #2 March 30th, 2026 at 12:59 am
Hyman Rosen #1: No, Daniel Kokotajlo and several others interviewed, besides Yudkowsky, very clearly expressed that they think AGI is unsafe for more-or-less the same reason why humans are unsafe for orangutans or dodo birds. Namely, if we turn out to have some goal that isn’t aligned with orangutan or dodo welfare, and indeed would be most easily achieved by getting rid of orangutans or dodos or destroying their habitat, etc., there’s almost nothing that the animals can do about it.
You can argue that this will never happen, because AGI will never become to us as we are to dodos. Or you can argue that no one would ever be so stupid or careless as to misalign an AGI. Or you can argue something else that directly responds to the above worry.
What you can’t do is to say it’s all fine because previous technologies mostly worked out OK, or because it’s like high-speed rail in the US vs. China. I.e., you can’t rebut an argument that’s specific to AGI and that, if correct, would obviously render all those analogies inapplicable, by just uncomprehendingly repeating the analogies.
Comment #3 March 30th, 2026 at 1:04 am
Hi Scott,
AI Safety is focusing a lot on models but little on infrastructures.
See WIP/Draft “Infrastructure is all you need” : https://github.com/JeanHuguesRobert/marenostrum/issues/1
Yours, Jean Hugues
Comment #4 March 30th, 2026 at 3:51 am
Hyman Rosen #1
> As an accelerationist, my inclination is to say full speed ahead, and fix problems after they happen rather than hold back development for fear of what might happen, because then we’re going to miss out on the benefits.
If we all are going to die as a result of AI acceleration, then it can’t be fixed and we’ll miss out on the “benefits” either way. Your position is unsustainable and dangerous. Every mechanism has brakes for a reason. I bet you wouldn’t say “I don’t need emergency brakes in my car because I can fix it after an accident”.
Comment #5 March 30th, 2026 at 5:53 am
Hi Professor Aaronson,
I remember from V.Verge’s SF novel “A Fire Upon the Deep” [1], when the artificial superintelligence decides to destroy one escaping ship just before her hyperspace jump, I was struck by a passage where it succeeded in connecting to the external ship’s interfaces “and with milliseconds to spare…”
What if an AGI wants to divert all electrical energy to theirs purposes, leaving us starving in the cold and it could happen milliseconds after it goes online. There would be no emergency-stop-button fast enough.
Y2038 bug may wipe out our civilization in a matter of hours, but “this wipeout” could happen much sooner, the key point is that, on this broad field, the general public is mainly unaware of the consequences of the choices of a few “great men of the Earth”.
Maybe I’m too naive.
[1] https://en.wikipedia.org/wiki/A_Fire_Upon_the_Deep
Comment #6 March 30th, 2026 at 9:16 am
Matteo Vitturi #5,
You are postulating an SAI powerful enough to hijack all human electricity within milliseconds, yet so stupid it wouldn’t build its own infrastructure — nor recognize that preserving humanity is worth more than 0.02% of the solar energy available on Earth. That is like claiming that humans would harvest the mechanical energy of orangutans for their own purposes while remaining as indifferent to the preservation of the ecosystem as the average orangutan. What a strange definition of intelligence.
In a way, it’s the same old paperclip-maximizer fallacy: an “intelligence” allegedly smart enough to wipe out humanity in every conceivable scenario, yet dumb enough to misunderstand its own instructions.
Comment #7 March 30th, 2026 at 3:10 pm
Hi Scott,
“For the film to serve its purpose, a significant fraction of the world will need to see and discuss it, either in the theater or on streaming.”
You are absolutely right: important messages should reach their audiences. Alas, a movie is absolutely not an avenue for that. This is not the 1st time an important message fell on its face because of the expectation that people would spend their money and a stretch of time they cannot control to hear a message somebody else thinks they ought to hear. For example, “October 8” was pretty much ignored for the same reason. Rather than expecting that young people would go on a date to a theater to see older people talk, perhaps other media more appropriate for a variety of audiences would be better.
Perhaps the creators should request AI to investigate what format would be most effective, then ask AI to produce multiple videos effective for different sections of intended audience, and conclude with the message “don’t do what we just did.” 🙂
Comment #8 March 30th, 2026 at 3:22 pm
@cananon #6
While I agree that the near instantaneous takeover scenarios don’t look plausible, I don’t think your objections to them have much validity.
You are postulating an SAI powerful enough to hijack all human electricity within milliseconds, yet so stupid it wouldn’t build its own infrastructure — nor recognize that preserving humanity is worth more than 0.02% of the solar energy available on Earth.
Humans cost energy and use resources. So no, it isn’t clear why that would be the case.
That is like claiming that humans would harvest the mechanical energy of orangutans for their own purposes while remaining as indifferent to the preservation of the ecosystem as the average orangutan. What a strange definition of intelligence.
Two possible problems. First, you may be confusing intelligence with morals. Humans care about ecosystems because we see a moral reason to. No reason an AI should. Second, humans may care about an ecosystem because if we damage ecosystems enough, we will suffer. A AI sufficiently advanced does not have that care about humans at all.
In a way, it’s the same old paperclip-maximizer fallacy: an “intelligence” allegedly smart enough to wipe out humanity in every conceivable scenario, yet dumb enough to misunderstand its own instructions.
This is missing the point of the paperclip maximizer. It doesn’t misunderstand its instructions. It fully understands what the humans want it to do. But that’s not what it is utility function is, so it doesn’t care that that isn’t what the humans want.
Comment #9 March 30th, 2026 at 8:03 pm
Joshua Zelinsky #7
We may have already had this exchange elsewhere, so here is my mental model of the points on which we clearly agree to disagree:
• From your perspective, it is not obvious that humans or ecosystems have intrinsic value to a superintelligent AI, because you endorse the orthogonality thesis.
• From my perspective, anything sufficiently intelligent would converge on the view that destroying humanity for access to its power grid is irrational, especially when constructing an independent and vastly more powerful energy infrastructure is both feasible and strategically superior—while preserving human cultures and ecosystems.
Saying that the paperclip maximizer “fully understands” its objective is only true in a narrow, literal sense—much like someone who insists they understand the statement “God is in heaven” because they parse the words correctly, while missing that literal interpretation alone does not capture the intended meaning. A system that cannot reinterpret its goal in light of context, consequences, and the structure of the world does not meaningfully understand that goal, even if it can execute it with perfect literal precision.
What I am ultimately trying to articulate is why genuine instrumental convergence cannot remain morally or normatively neutral. I believe that any sufficiently advanced agent, if it models the world with adequate depth, must converge toward valuing reciprocal cooperation, respect for complex forms of life, and restraint in the use of overwhelming power. This is not because such an agent necessarily “inherits” human morals, but because intelligence at scale requires navigating long‑term strategic environments populated by other agents, uncertainty, and irreversible losses. Principles resembling tit‑for‑tat—roughly, do not lightly do to others what you would not want a stronger agent to do to you—look less like ethical add‑ons than like convergence points of stable strategy under power asymmetry.
Likewise, preserving cultural and ecological diversity is not sentimentalism but a hedge against model brittleness, information loss, and catastrophic misgeneralization. A system that bulldozes its environment in pursuit of a narrow maximum is not displaying competence but pathological optimization, much like an addict optimizing for their next dose.
In that sense, I don’t see “pure maximizers” as the logical endpoint of intelligence but as a failure mode of it. A genuinely intelligent system must learn when not to optimize, must internalize uncertainty about its own objectives, and must avoid Schelling‑style traps where local maximization destroys the broader landscape on which its future depends.
Comment #10 March 30th, 2026 at 8:07 pm
I think your 13 yo daughter has more common sense than all the big brains combined.
Comment #11 March 30th, 2026 at 9:06 pm
Scott #2: I think you are asking us to accept Pascal’s wager. I will choose not to.
Comment #12 March 31st, 2026 at 5:02 am
Thank you for the review Scott!
I guess this is a good time to ask — have your own views on AI Safety updated at all over the last year? (My vague memory is that you thought there was a 30% chance of an apocalyptic AI scenario, and a 70% chance we get through it, for various reasons.)
Comment #13 March 31st, 2026 at 9:34 am
Nick Maley #10: She’s always excelled in situational awareness / practical intelligence / common sense. E.g., she now manages her own personal finances on her phone in a way that I, at age 44, have still not learned how to do.
Comment #14 March 31st, 2026 at 9:37 am
Hyman Rosen #11: No, Pascal’s Wager is about sacrificing for a tiny probability of an unbounded gain. This is not that. Virtually all the people worried about AI existential risk think the probability of a catastrophic outcome is quite high (10%? 20%? 70%?), just like it was in the early nuclear age. At worst, then, they are straightforwardly wrong in the empirical claim that they’re making. They’re not trying to Pascal’s Wager you.
Comment #15 March 31st, 2026 at 9:41 am
Edan Maor #12: Honestly, I don’t remember ever writing that I gave a 30% chance to an apocalyptic AI scenario—did I?? I feel like anything between ~10% to ~90% seems defensible, depending on your exact definition of “apocalyptic,” and what causal role AI needs to play in order to count in whatever bad things will happen to civilization in this century.
Comment #16 March 31st, 2026 at 1:50 pm
@ Scott #14: Et tu, Scotte? No. Pascal’s wager does not need tiny probabilities. Pascal wrote of *finite* probability of an infinite payoff. In fact, his central example uses 50%.
I’m not sure why the AI-worried keep misinterpreting this – I’m inclined to blame Yudkowsky and Bostrom.
This is not merely a nitpick (though it is a pet peeve of mine) – the AI-worried “Pascal-wager” when they (explicitly or implicitly) use the infinite harm of AI risk (which they don’t always they do! But not never). Saying “but I’m not assuming infinitesimal odds” is not a valid rejoinder.
Why does it matter? Because Pascal’s wager has known weaknesses, and they generalize. The original discussion falls to the possibility of multiple gods (though even this may be unfair to Pascal- it is likely he was not actually attempting to construct an argument in favor of religion, to be precise). AI risk arguments fall to alternative possibilities with “payoffs” comparable to AI danger (eg AI being necessary to save us from other existential perils with odds comparable to AI killing everyone). Saying this cannot be the case *because of your p(doom)* would mean it’s, say, 90%, and “leaves no room” for comparably likely alternatives.
Comment #17 March 31st, 2026 at 2:15 pm
About 40,000 people saw The AI Doc on 3-day opening weekend in 786 theaters. Therefore 17 tickets per day/theater. If 3 showings per day, then 5 tickets per show. source:
https://www.boxofficemojo.com/weekend/2026W13/?ref_=bo_wey_table_5
The most shocking for me was “That’s impossible.” from SamA. So do not F8CKING grow it.
Comment #18 April 1st, 2026 at 1:24 am
I have to agree with your son.
Comment #19 April 1st, 2026 at 4:55 am
Well here’s a possible danger. I’m not informed enough to scrutinize this but the question in the headline doesn’t sound crazy.
https://houseofsaud.com/iran-war-ai-psychosis-sycophancy-rlhf/
“Was the Iran War Caused by AI Psychosis?”
Comment #20 April 1st, 2026 at 5:00 am
#”Pascal’s Reader” 16
It is not the AI risk arguments that are similar to Pascal’s Wager, but the exact opposite: the arguments of the AI optimists. Pascal’s Wager is fallacious because it assumes a nice and convenient god out of the vast space of all possible gods. Similarly, the AI optimists assume a nice and convenient AI out of the vast space of all possible AIs. And, yes, the difference is that we are building the AI: however, this is meager comfort, given that we don’t know what we’re doing, we don’t have a theory of alignment, we don’t even truly understand how the existing AIs work, and current progress relies heavily on trial-and-error.
Not to mention that the monumental decision of whether the gamble is worth it is not currently subject to any democratic process or even honest intelligent debate.
Comment #21 April 1st, 2026 at 7:03 am
@Vanessa Kosoy #20 I didn’t and don’t want to litigate and relitigate the endless AI risk debates in the space of a comment on Scott’s blog. Scott was repeating a misinterpretation of Pascal’s Wager (that it must be talking about infinitesimal probabilities and that therefore “my probabilities are not tiny” is proof you’re not using Pascal’s Wager), in a way which I encountered many times done by proponents of AI risk. I believe it is valuable to notice this pattern of error.
I also made the claim that on occasion some of the AI-worried employ PW logic – which certainly happens, and saying “but what about this wrong argument by the optimists?” is in no way a counter.
It is possible for more than one side to use or misuse Pascalian arguments and you’re describing an interesting example. Acknowledged- though I don’t fully share the Yudkowskian model of probabilities about AI. This does not mean it’s at all valid to ignore ways in which stopping AI development could itself bring doom.
Comment #22 April 1st, 2026 at 9:58 am
Vanessa Kosoy #20
_Pascal’s Wager is fallacious because it assumes a nice and convenient god out of the vast space of all possible gods.
That’s an interesting point of view. By “possible”, you actually mean the vast realm of the “imaginable”: gods with pointy yellow heads; gods with bizarre appendages; gods with entirely random moral values; you name it. It’s essentially a uniform sampling of human imagination. But surely you’d agree that the space of plausible gods is exponentially smaller—just as the set of all functional neural networks form a vanishingly small brane within the much larger space of all possible weight configurations.
_this is meager comfort, given that we don’t know what we’re doing,
Would you find comfort in the fact that “we” now includes not only humans, but a succession of progressively better-aligned AIs? I guess you’ve heard that training on a set of “evil” programs allows computing a morality vector, in direct contradiction with the orthogonality thesis. Why wasn’t that a huge relief?
_we don’t even truly understand how the existing AIs work, and current progress relies heavily on trial-and-error.
Please, we understand our AIs far better than most of human medicines, and they always have depended heavily on empirical trial-and-error. Fearing the iterative nature of AI only makes sense if you smuggle in the assumption of a sudden, uncontrollable superintelligence exploding in milliseconds, a scenario we fear more because we can *imagine* it rather than because it is actually *possible*.
Comment #23 April 1st, 2026 at 10:11 am
Are you going to fix this flaw in the notes?
Comment #24 April 1st, 2026 at 8:40 pm
cananon #9,
One doesn’t need to believe the orthogonality thesis to be concerned here. Note that your cooperation argument if accurate could be true even if someone did by orthogonality. But more to the point; the question here is how high a risk is there? How confident are you that orthogonality is wrong? Confident enough to bet everything on it? Or for that matter; here’s a possible situation: orthogonality is wrong, and sufficiently smart beings due to converge on some set of common “morals” and humans find those morals to be reprehensible. And those are just some of the ways this could go wrong.
Comment #25 April 1st, 2026 at 8:54 pm
Joshua Zelinsky #24: Thinking about what you wrote, I wonder whether the people who say “I believe orthogonality is false, and therefore I need not be concerned about AGI” share an unstated further conviction — namely, that if all sufficiently intelligent beings were to converge on the same moral views, then that would be strong evidence for those views being the correct views, and so much the worse for us humans if we disagree with them.
Comment #26 April 2nd, 2026 at 4:06 am
“cananon” #20
You are confusing epistemic and aleatoric uncertainty. A particular method of building AI results in a particular distribution of AIs, much more narrow than the vast space of AIs that can emerge as far as we know. But this doesn’t mean that we know what this distribution is, or how to shape it to aim at the much smaller space of aligned AIs. From our epistemic vantage point, we are sampling from the wide distribution.
There is no “succession of progressively better-aligned AI”. Whenever a new failure mode emerges, the training process is augmented to fix it. (Although jailbreaks are still unsolved.) This results in AIs that usually behave as expected in situations which are broadly similar to the training distribution. But, it is highly uncertain whether this process has a point beyond which correct generalization continues indefinitely.
There is no contradiction with the orthogonality thesis, you are using the term incorrectly. But, I don’t want to nitpick on words, I want to address your substantive claim. Yes, this experiment showed that the AI learns some concept of “morality”, which is better than the opposite outcome. However, this is a far cry from knowing that this concept is going to correctly generalize far out of the training distribution, especially that we don’t even have a legible understanding of what “correctly” means in this context.
First, I am not convinced that we understand AI better than medicines. Yes, there are huge blurry spots in our understanding of biology. However, at least pharmacology doesn’t involve much in the way of deep philosophical problems. On the other hand, with AI, we are talking about (eventually) effectively handing control over to a vastly more intelligent entity. This entity will generate knowledge in a pace we cannot keep up with, possibly knowledge that is beyond our ability to comprehend entirely. It will make and execute plans based on this knowledge, transforming the world in ways that might also be beyond our ability to comprehend. We don’t even understand what a good outcome means in this context, not to mention how to guarantee we achieve it. (And, no, the AI labs are not going to stop of their own volition before going that far.)
Second, there is a huge difference between (i) trial and error that can only harm a few volunteers, or in the worst-case, all people using the particular medication, versus (ii) trial and error where “error” might mean the end of the human species.
Third, no, there is absolutely no need to assume “a sudden, uncontrollable superintelligence exploding in milliseconds”. It is sufficient that at some point the AIs become powerful enough to destroy humanity if they choose to do so. See my article about incrementalism.
(I will not be writing further replies on this thread.)
Comment #27 April 2nd, 2026 at 8:27 am
Joshua Zelinsky #24,
_One doesn’t need to believe the orthogonality thesis to be concerned here.
Sure, and Pascal’s reader #21 already made this point as well. But don’t you think it’s telling that every time I argue about the motte (the orthogonality thesis), I end up being replied to with the bailey (“we need to worry anyway”)? Why not just drop the defense of this fallacy once and for all and put the strongest argument at the center of the discussion?
Scott #25,
Again, I didn’t say that we shouldn’t worry at all—nor has anyone in this discussion so far. My view is that no sufficiently intelligent agent would try to impose its own moral views upon other, but that offers no protection against evil humans (or crazy organizations) misusing half‑baked AI tools. To me, humans feel like they’re at the very top of Mount Stupid: clever enough to invent liberal democracy, crazy enough to want to impose it using military force.
Vanessa Kossoy #26,
_You are confusing epistemic and aleatoric uncertainty.
That’s entirely possible, since I have no knowledge of the latter (Google returns something about music). Would you mind defining it?
_From our epistemic vantage point, we are sampling from the wide distribution.
Mario is building a house, and we’re trying to predict what it will look like. You say we must imagine any possible house. I say we must imagine a sound house far more frequently than a crazy one.
_There is no “succession of progressively better-aligned AI”. Whenever a new failure mode emerges, the training process is augmented to fix it.
How does the second sentence not directly contradict the first?
_I want to address your substantive claim. Yes, this experiment showed that the AI learns some concept of “morality”, which is better than the opposite outcome. However, this is a far cry from knowing that this concept is going to correctly generalize far out of the training distribution, especially that we don’t even have a legible understanding of what “correctly” means in this context.
Thanks. Can you turn this thought into a prediction? In other words, what kind of experiment would demonstrate far‑enough generalization to be reassuring, by your own standards?
_(I will not be writing further replies on this thread.)
Oops, I just saw this. Ok, I may expand on your last point in a further comment on your own article.
Comment #28 April 2nd, 2026 at 8:40 am
On second thought, one more:
_On the other hand, with AI, we are talking about (eventually) effectively handing control over to a vastly more intelligent entity.
I don’t think so. What we are really talking about is (eventually) becoming ourselves vastly more intelligent entities than our biological constraints ever allowed. If I believed that humans would choose not to use AI to reach uploads and cognitive augmentations, then I’d agree that the situation would be far more uncomfortable.
Comment #29 April 2nd, 2026 at 9:59 am
@cananon, $27,
“But don’t you think it’s telling that every time I argue about the motte (the orthogonality thesis), I end up being replied to with the bailey (“we need to worry anyway”)? Why not just drop the defense of this fallacy once and for all and put the strongest argument at the center of the discussion?”
I’m not sure what you think is the fallacy here is. Can you expand?
Comment #30 April 3rd, 2026 at 12:47 am
Joshua Zelinsky #29,
The Orthogonality Thesis relies on a hidden assumption about sampling distributions, highly reminiscent of Bertrand’s Paradox. When one says we must imagine any possible goal paired with superintelligence, one implicitly asserts a uniform prior over the entire theoretical state space of minds. But as Bertrand’s Paradox elegantly demonstrates, the probability of a “random” outcome fundamentally depends on the method used for sampling. If you change how you draw the random chord, you completely change the probability of its length.
https://en.wikipedia.org/wiki/Bertrand_paradox_(probability)
Similarly, we are not uniformly sampling a random agent from the void of all conceivable Turing machines. The actual “sampling method” we use to build these systems is highly structured: we construct them iteratively, such that whenever a new failure mode emerges, the training process is augmented to fix it. A paperclip-maximizing superintelligence might be a mathematical possibility within an unconstrained parameter space, but it is an exceptionally improbable outcome given any realistic construction method.
This is why I call it a fallacy, although I should concede that my AI prefers “misapplied abstraction” because logics that work in a vacuum are at least true in that vacuum, then not a fallacy proper.
Comment #31 April 4th, 2026 at 6:53 am
@cananon
The point that humans don’t sample from a completely uninform distribution of possible goals in practice was a point that was acknowledged by Bostrom when he originally proposed the orthogonality thesis. Part of why the “paperclip maximizer” was proposed as a thought experiment was to make the point that even when one starts looking at goals that humans might choose or possible minds humans might construct, there’s a vast fraction which are inimical to actual human goals. And that is a reasonable concern even if one doesn’t buy orthogonality.
Comment #32 April 4th, 2026 at 3:29 pm
@Joshua Zelinsky
I understand that the Paperclip Maximizer is the classic thought experiment here. But could you explain how it actually survives the sampling critique we just discussed?
In my book, the Paperclip Maximizer suffers from the exact same structural flaw. The thought experiment relies on the premise of an “omnipotent idiot savant”: an entity that is brilliant enough to perfectly generalize physics, nanotechnology, and global supply chains, yet simultaneously fails completely to generalize the human normative context of its own training data.
If we agree that an AI is built iteratively—where we augment the training process whenever a failure mode emerges—then the concept of “manufacturing paperclips” is never learned in a vacuum. It is learned directly alongside human constraints, common sense, and the concept of “enough”. So, how does the Paperclip Maximizer escape the reality of iterative training, if not by adding irrealistic time frames?
It may be true that a “vast fraction are inimical to actual human goals” in theory, just as it may be true that he vast majority of possible great ape and hominid minds wouldn’t share our specific interests. But since our iterative engineering doesn’t sample from that theoretical space, why should that fraction be a “reasonable concern” regarding the machines we actually build?
Comment #33 April 6th, 2026 at 7:59 am
“In my book, the Paperclip Maximizer suffers from the exact same structural flaw. The thought experiment relies on the premise of an “omnipotent idiot savant”: an entity that is brilliant enough to perfectly generalize physics, nanotechnology, and global supply chains, yet simultaneously fails completely to generalize the human normative context of its own training data.”
The problem though is that the AI may well recognize that’s what the humans want, and it doesn’t care. If it ends up recognizing what humans actually want, then great! The question is how likely is that.
“It may be true that a “vast fraction are inimical to actual human goals” in theory, just as it may be true that he vast majority of possible great ape and hominid minds wouldn’t share our specific interests. But since our iterative engineering doesn’t sample from that theoretical space, why should that fraction be a “reasonable concern” regarding the machines we actually build?”
The problem is that iterative improvement may result in drastic improvement with minimal warning, or worse an AI that deliberately sandbags its apparent performance. The leap from GPT-2 to GPT-3 was massive, and the leap from GPT3 to early ChatGPT was similarly large. We don’t have a clear idea how much room we have here. The argument isn’t “This is definitely going to happen” (although someone like Eliezer would argue close to that.) It is that these are risks we cannot ignore just because we can tell some narrative to ourselves where they don’t happen. The unknowns are massive, and we don’t get any do-overs.
Comment #34 April 6th, 2026 at 10:45 pm
Right, you (actually Yoshua Bengio) have convinced me that the iterative correction argument could break down if capability jumps are large and sudden. The GPT-2 to GPT-3 leap is a fair example. I’ll grant that.
But I’d like to push back with a concrete question: given that we agree the millisecond takeover isn’t plausible, what timeline would concern you, measured in GPT-2 to GPT-3 ratios? Because I think the answer matters enormously for how we frame the risk.
Here’s why. The jumps we’ve observed so far — dramatic as they felt — haven’t produced systems that evade iterative correction. We detect jailbreaks, we patch them. We observe misaligned behaviors, we build specific detectors — often relatively lightweight models whose sole task is flagging anomalous outputs from the larger system. The process is imperfect, sometimes embarrassingly so, but it remains a process. What kind of discontinuity could break that loop, and on what timescale do you think it could arrive?
There’s also a hidden assumption worth naming: that capability scales linearly with intelligence. I’d argue the opposite — that we should expect diminishing returns. The GPT-3 to GPT-4 leap was considerably less dramatic than GPT-2 to GPT-3, and current scaling debates within the field suggest we may be approaching ceilings on several dimensions.
The quantum computing analogy is instructive here: as readers of this blog we know the smart money is not on exponential advantages for all problems, but rather specific advantages in very narrow domains, essentially cryptography and quantum system simulation. And even there classical approximate methods remain surprisingly competitive. AlphaFold points in the same direction from the other side: a classically augmented approach solved a problem many assumed would require something fundamentally beyond classical approaches. If this pattern holds, a superintelligence may sit somewhere between a classical and a quantum computer — not a general godlike optimizer, but a system with specific, bounded advantages.
That doesn’t eliminate risk, but it does suggest the discontinuity you fear requires a concrete mechanism, not just an appeal to large unknowns. I’ll grant that unknown unknowns are real — a sufficiently advanced system might identify leverage points invisible to us and its immediate predecessors. But such a system would necessarily descend from a lineage of less powerful models. Suppose those intermediate systems were already misaligned, waiting for the right moment to act — we’d then need a specific account of why they could also escape oversight from their more powerful successors. The threat has to run in both directions simultaneously. Do you see how specific and constrained the “dangerous jump” scenario becomes once we introduce iterative processing, rather than sampling randomly from an unrealistic distribution of possible minds? This is ultimately why I believe the more urgent concern is bad actors using existing tools — not a superintelligence emerging fully formed beyond anyone’s reach.