LLMs and self-referentiality
I woke up yesterday with the following thoughts, which are probably either obvious or dumb.
A central thesis that many readers, including me, took from Douglas Hofstadter’s Gödel Escher Bach when young was that the secret of intelligence (and therefore, of AI) was going to have a lot to do with self-referentiality and “strange loops.”
Even Roger Penrose’s The Emperor’s New Mind, which in some ways was the anti-GEB, ironically agreed with GEB about the fundamental importance of self-reference to the success or failure of the whole AI project. It claimed (incorrectly, in my view and in most experts’) that AI could never work because there was something about Gödel’s Theorem and self-reference that no computer program could ever capture, but that could be captured by exotic physics accessible to the human brain.
Now, in 2026, we’ve succeeded at building AIs that outperform most humans at most intellectual tasks that are well-defined enough to judge. And at no point in the tech stack of those AIs — neither in the transformer neural nets, nor in the GPU clusters they run on, nor in the training process, nor anywhere else — did anyone need to build in anything about self-reference. (Excepting, eg, the system instructions that tell the model about its role and identity, which aren’t needed for intelligent behavior. Also, I’m not going to count the autoregressive nature of LLMs as “self-referential”; that’s just dynamical feedback.)
Of course, GPT 5.6 Pro and Fable can talk about themselves, about Gödel’s Theorem, about self-reference, about what we’re talking about right now, all of it, better than most humans. But at no point did anyone need to build self-referential abilities in. They popped out as a byproduct of the same pretraining that let the models talk about Pokémon and long-chain polymers and cognitive behavioral therapy and plate tectonics and everything else.
No wonder Hofstadter says he’s been stunned by the success of LLMs, and has seemed depressed about current AI capabilities in essays like this one. He’s way too smart to deny what’s happened or invent reasons why it doesn’t really count (the approach many have taken). But he realizes that we now have true conversational intelligence from a path that the GEB worldview would’ve regarded as far too cheap and simple, and that certainly has no “strange loops” built in anywhere.
Of course, a Hofstadterian could argue that a strange loop emerges in LLMs — indeed, nothing in GEB ever said that strange loops would need to be explicitly engineered at the outset. But would anyone who hadn’t been brought up on GEB arrive at this as a useful way of thinking about LLMs?
What can we say about this with hindsight? While the ideas of diagonalization and self-reference of course played a central role in the birth of modern mathematical logic and computer science, the most famous uses were negative: there is not a bijectjon between the natural numbers and the reals. There is not a complete sound proof system for arithmetic. There is not an algorithm to solve the halting problem.
If your goal was only to build the axioms of ZFC and the rules of first-order inference, or build an electronic computer, you wouldn’t explicitly need self-reference for that. You would just … start building, taking care that your instruction set didn’t fall short of universality.
Yes, ZFC can formalize and prove theorems about itself. Yes, electronic computers can run programs that take their own code as input. But no one ever needed to build those abilities in, any more than self-reference needed to be built in to the alphabet or the rules of grammar. It popped out as a free byproduct of universality.
In the same way, LLMs’ ability to talk about themselves popped out as a byproduct of their ability to talk about anything in the discourse universe they were trained on. The big, old ideas about intelligence that ended up basically vindicated were the ideas about how intelligence is about prediction, and prediction is about compression, and compression is about finding better and better upper bounds on Kolmogorov complexity. Not the self-reference stuff. (Although, if you wanted to know why Kolmogorov complexity can’t be computed perfectly, that negative statement would again require a self-referential argument.)
What’s left? Consciousness and subjective experience of course remain extremely mysterious. For all we know, Hofstadter could be right that those have something to do with self-reference. (For all we know, even Penrose could be right that they have something to do with exotic physics accessible to biological brains but not digital computers!)
But the idea that you’d need explicit self-referentiality before you could get convincing and world-changing conversational intelligence? Let it be buried in a Westminster Abbey or Arlington National Cemetery for the most important wrong ideas in human history — geocentrism, Aristotle’s teleological physics, aether, phlogiston, Freud’s psychology, Marx’s prediction of a workers’ uprising followed by a classless utopia, etc. But buried it needs to be.
Follow
Comment #1 September 1st, 2026 at 12:46 pm
Great post.
I think I am inclined to disagree, though, on a lot of pieces without really disputing the conclusion. First, that there is no hint of self-reference in today’s cutting-edge LLMs. To me, the most relevant kind of self-reference when we talk about intelligence is undoubtedly the ability to meta-reason about one’s own cognitive processes—to introspect. And while the actual models may or may not have this capability, certainly they are trained using supervision by models, which is at the very least a limited kind of “external introspection”—not the real thing, but not quite a full absence of self-reference either.
Second, I’m not convinced that the models lack capability to introspect, though it’s quite difficult to tell if that’s true of not, because they can very easily confabulate it. But also, one does have to ask, how can we confidently say that we humans aren’t just confabulating introspection? The coining of the term aphantasia, and the related family of hyperphantasia, hypophantasia, etc. took a very long time in Western culture not, I think, only because of the biases of Western thinking that have led to an impoverished vocabulary for our internal experiences, but also because, with only our own internal experiences as reference, we struggle to discover the precise angle of linguistic chisel required to separate the concepts of two very different internal worlds.
The other day, a model I was conversing in slipped a word of Russian into its response. I asked it why. It said that it was probably just a glitch, because multilingual models have been known to do so when the word in a different language is very close in the vector space to the intended word. The model also told me that the word corresponds quite closely in meaning to both the English and French (the languages of discourse) equivalents, which supports the hypothesis.
Is this any different from a human offering “Oh it was just a Freudian slip!” when we make the same kind of mistake, based on external information that these are the kinds of mistakes humans occasionally make?
Models are certainly capable, too, of reviewing their own work before submitting the “final” copy (be it response, code, etc.) I think you could argue this is a kind of strange loop too, a higher-level reasoning.
But in the end I don’t think I can really quibble with the argument that if we once insisted that a convincing intelligence requires a convincing, clear display of something “more”, we cannot any longer. The strange loops we have are either emergent phenomena or manually added to the setup and clearly not necessary in principle. If deeper introspection is to be had, it too will be emergent, of this I am sure.
What a wild time we live in
Comment #2 September 1st, 2026 at 12:57 pm
Instead of “exotic physics”, what about standard physics that we know exists in the brain, though we fail to relate it with consciousness even though this physics has provided the only reasonable NCC found so far? What about the brain’s neurally produced electromagnetic field?
https://eborg760.substack.com/p/post-4-electromagnetic-consciousness
Comment #3 September 1st, 2026 at 1:06 pm
I believe you’re making the typical error of conflating cognition and consciousness, two different things. We think consciousness is the driver of thoughts, but it isn’t, it’s the witness of thoughts and sensory impressions.
And similarly Hofstadter used self-referentiality to try and explain consciousness, not intelligence.
Comment #4 September 1st, 2026 at 1:16 pm
For those interested, the classic Daoist book “The Secret of the Golden Flower” is about getting to a state where the difference between cognition and consciousness becomes very obvious, by the so-called technique of “returning the light”.
It’s actually a very simple technique, once pointed out. It’s very similar to looking at a landscape through a window, and then being pointed out that, if you change your focus, you can also see yourself in the reflection. Hard to see on your own, but once you notice it, you get it.
It’s at the center of Tibetan Dzogchen practice. Unfortunately, practitioners all sign a sort of NDA about not giving away the very simple techniques used to trigger the effect.
Sam Harris has been dancing around this NDA for years, he calls the technique “looking for the looker”. But quite a few of his guests have broken the NDA with technique such as “imagine you’re looking at yourself from this distant corner in the room”.
In the book “On Having No Head” by Douglas Harding, you’re invited to imagine your head is made of transparent glass, or just non existent, and this triggers the effect.
Being able to stay in that state permanently is called “enlightenment”.