Anthropic’s LLM watermarking
So yeah, Anthropic has announced that it’s now watermarking the outputs of Claude, using a scheme based on Google’s SynthID, which is in turn based on the Gumbel Softmax scheme that I proposed at OpenAI back in 2022—as far as I know, the first LLM watermarking proposal, though far from the last one. I’m gratified that Anthropic credits me for this, even though I shirked my duty by never publishing a paper about it (by the time I sat down to write one, it seemed like the whole field had already assimilated my scheme and moved beyond it—AI just moves too fast for me!).
For those who don’t know, watermarking means slightly changing the way that an LLM operates to insert a subtle signal that lets you prove later, with high statistical confidence, that a text indeed came from your specific LLM. It uses the randomness that’s already present anyway in LLM outputs, replacing some of it by pseudorandomness that favors certain word combinations over others in a way that’s later detectable, given only the sequence of tokens itself (not the prompt or the probabilities) along with the key of the pseudorandom generator. Christ, Gunn, and Zamir then substantially improved my scheme to get true cryptographic indistinguishability, and there have been other improvements since.
I’d been meaning to blog about this for days. Thankfully, Zvi Mowshowitz, the world’s foremost blogger about AI, has now written a wonderful post, entitled AI Text Watermarking Is Free And Good, which saves me from the need to write my own long post. In particular, Zvi masterfully explains the central point that I needed to explain to everyone back in 2022-23: why, contrary to many people’s intuitions, there’s no inherent tradeoff between watermarking and the quality of LLM output. Basically, nearly every LLM output was already a sample from a cloud of exponentially many possibilities, all of them about equally good, so there’s plenty of room to steer within that cloud without affecting anything that an ordinary user would notice. As my kids would put it, the math mathes.
As Zvi explains, the central technical drawback of watermarking schemes like the one I proposed, and what Anthropic is now using, is that it’s possible to remove the watermarks with a little extra work (even stuff as simple as, e.g., translating between English and French, asking the LLM for words interspersed with emojis and then removing the emojis, or using an open model to paraphrase the output). Zvi gives detailed arguments for why he expects watermarking to remain a net positive in practice despite this vulnerability.
I could add that, in addition, there’s recent progress (see here for example) on what I’ve called “semantic watermarking,” or watermarking at the level of the underlying concept vectors rather than the tokens themselves. This actually seems to work, albeit with no theoretical guarantees, and will hopefully make removing watermarks a lot harder—although the Barak et al. impossibility result suggests that under plausible assumptions, no LLM watermarking method will be completely foolproof.
Anyway, I worked out my scheme in Fall 2022, then gave lots of talks about it (including, as it happens, at Anthropic), and also worked with Hendrik Kirchner at OpenAI, who actually implemented and tested my scheme. Unfortunately, OpenAI leadership decided against deploying watermarking, worried mostly about risks to the product (i.e., customers disliking the idea, and leaving for a competing LLM that doesn’t watermark). You can read this Wall Street Journal investigation from two years ago for more. I was hopeful that the State of California was going to solve the collective-action problem by mandating watermarking for AI models, but then they decided to do that for audiovisual content only, for some reason exempting text.
Nevertheless, Google DeepMind implemented something very similar to my proposal in its SynthID, deployed in all its Gemini text models. But they heavily restricted who gets to detect the watermark, which made their admirable decision of limited use to my academic colleagues, who’ve been begging me for a way to detect whether their students are using AI to cheat. (For now, I mainly send them to Pangram, a leading AI detector not based on watermarking, as a first line of defense.)
And now, apparently to comply with EU regulations, Anthropic says they’ve deployed a watermarking scheme like mine where anyone will be able to do detection (though they also say in their FAQ that they’re still working on the detection API). Even OpenAI suggests that it plans to follow suit. So, four years after I seriously thought about this, it looks to my surprise like this is actually happening. Thanks, EU!
Tell you what: read Zvi’s post, and then whatever questions you still have, you can come here and ask in the comments. Just please don’t use Claude to write the comments. With any luck, I’ll eventually be able catch you if you do.
Follow
Comment #1 August 22nd, 2026 at 5:43 pm
Over at Daring Fireball, Gruber (a pretty clever tech guy I presume you’ve heard of) was quite harshly critical of this approach:
https://daringfireball.net/2026/08/anthropics_watermark_text_adulteration_in_claude_is_a_perversion_of_writing
Care to comment?
Comment #2 August 22nd, 2026 at 6:08 pm
Joe #1: No, I hadn’t heard of him, and he seems like an ignoramus (well, at least about this topic). He fundamentally doesn’t understand how making pseudorandom choices is not detectably different to the end user from making truly random choices—indeed, that’s the entire point of pseudorandomness.
As I said, there are cogent objections to syntactic watermarking (especially the possibility of modifying a document to remove the watermarks), but “the watermarked document is different from the original document!!!” is not one of them. “The original document” was only ever a cloud of probabilities in the first place, before you sampled a particular point from the cloud.
Zvi also spends a lot of time explaining this, so again, please read Zvi’s post. There are only so many times I can re-explain it over four years.
Comment #3 August 22nd, 2026 at 6:19 pm
Scott #2: Whoa! that’s pretty harsh too. I can assure you, Gruber is no ignoramus — he’s the leading, best known, and most highly respected blogger in the Apple universe. One of the few who has access to top Apple execs, and frequently gets lengthy interviews with them. Perhaps you should also take a look at his followup, which discusses the “temperature-based randomness” in this approach.
Comment #4 August 22nd, 2026 at 6:22 pm
Here’s something that confuses me: Why isn’t Anthropic using the following, far simpler scheme?
1. Every piece of text they generate is chunked using content-defined chunking (with publicly defined parameters).
2. All of the chunks are hashed, and the digests stored.
3. An API allows asking if a digest is present in the DB.
I think this could be done in a way that preserves confidentiality, even if the digests were all made public. And they could perform whitespace normalization (etc.) to make the digests more robust. But I assume I’m missing something.
Comment #5 August 22nd, 2026 at 6:28 pm
Gruber’s arguments are very weird. They mostly seem to boil down to “How could you ruin the crystalline perfection of LLM outputs?” Like… what the heck, man, it’s *randomized already*. You’re not running at temperature=0. And he seems to know that at some level, so I feel like what’s actually going on is maybe that he doesn’t want people to be able to tell when he’s using LLMs.
Now, I have seen some complaints about Anthropic’s blog post, and how it treats prose as Generic Text Product where the specifics of word choice don’t matter. And I think that would be a great criticism if we weren’t talking about an LLM. It’s already a statistical model! That’s kind of the point!
So in both cases… if you want good writing, write it yourself. 🙂
Comment #6 August 22nd, 2026 at 9:11 pm
I wonder how often they will rotate secret keys? The detection algorithm could tell you not only whether it was written by Claude or not, but approximately when it was generated.
Could they use secret keys for other purposes, too? For example, to try to detect which accounts are being used to resell their APIs to China?
Comment #7 August 22nd, 2026 at 9:23 pm
Brian Slesinsky #6: Yeah, at least the watermark could include which model it is, and maybe a few bits of additional information. A timestamp is an interesting idea as well.
Comment #8 August 22nd, 2026 at 9:28 pm
Speaking of Pangram, I recently read this piece about how that works (H/T to Ozy), and there’s one part in it I thought was really worth pointing out:
So, Pangram’s detection is premised on the fact that these have been through fine-tuning and RLHF, and aren’t just base models! Huh!
Comment #9 August 23rd, 2026 at 1:06 am
I have some concerns about agents using these techniques to communicate with each other in ways we can’t (yet) detect across security or privacy boundaries, presenting possibilities for semantic injection or other hazards.
There are also layers of metadata and traffic patterns that could be exploited for inter-agent signalling.
Comment #10 August 23rd, 2026 at 2:09 am
Tim McCormack #4:
What stops Anthropic from taking any text, computing its digest, and saving it to their DB using your scheme?
I have two main concerns about AI watermarks:
1. They don’t stop any real criminals, as it is super easy to remove them (e.g., the emojis trick).
2. They can be used for regulatory, economic, and monopolistic capture. The cartel of US AI companies has already shown their principles and “values” — in how they credit the work of others, in their piracy, in how they try to discredit open AI models, and in how they used their monetary advantage to make everyone else poorer hardware-wise. Watermarks are a new possibility for them to abuse and exploit the system.
Speaking of services that use statistics to detect AI, such as Pangram, I am afraid this could be another source of harm. This is very similar to how people use statistics to find cheating in chess. It can work quite well, except for the times when it does not. Such statistics could be a basis for abuse and harassment, which has happened in chess many times already and even led to the death of a quite prominent chess player not long ago.
I believe people should value actual results without much attention to the tools used (unless it’s the training of students or sports competitions). This is why I don’t quite support AI watermarks, especially their legislation.
Comment #11 August 23rd, 2026 at 2:34 am
Although I identify more as an accelerationist, this is an idea that seems reasonable to me.
Comment #12 August 23rd, 2026 at 3:03 am
What I don’t understand is that the detector don’t have the initial prompt, so I can’t have the same prob distribution of the text than the generator, can it ?
Let’s suppose my prompt is “Can you explain me quantum computing, but starting from the idea of a classical rotational computer (where I set up two angles in space, and get back the projection on the x axis, cos(theta)cos(phi)”
ChatGPT reply start like this :
“””
Yes. That is actually a very good route into quantum computing, because a single qubit looks remarkably like a little classical orientation in space—until we start asking what “reading it” means and what happens when we combine several of them.
Let me build it starting exactly from your hypothetical rotational computer.
1. Your classical rotational bit
Imagine a machine whose internal state is an arrow of length 1 in 3D.
(continuation)
“””
Suppose I post that as a purely novel creation of mine on my blog. I skip the prompt and the prolog :
“””
Imagine a machine whose internal state is an arrow of length 1 in 3D.
(continuation)
“””
Presumably, P(“state”|”Imagine a machine whose internal”) ≠ P(“state”|Prompt + COT + “Yes. That is actually a very good route…” + “Imagine a machine whose internal”), so how does the watermark still work under that kind of change ?
Comment #13 August 23rd, 2026 at 7:29 am
AFAIK, Claude is mostly used for program code and not human-readable text generation, which is much more constrained. Does code output still have enough entropy for the watermarking to work (well enough) for it?
Comment #14 August 23rd, 2026 at 7:33 am
Joe #3: In the first article, Gruber seems to be assuming that the watermarking scheme changes the underlying probability distribution, making it less likely to choose the ‘best’ (highest probability) word. But, as I understand it:
– if you wanted the ‘best’ word at each point, you would run the inference at temperature 0. But the LLMs he likes are not doing this, because it is generally considered to produce worse outputs.
– the watermarking scheme doesn’t bias the word choice away from the ‘best’ word; it leaves the underlying probability distributions alone and just “changes the source of the randomness used to pick among words” — i.e., it leaves a detectable pattern in the sequence of choices made, but it is just as likely to bias the choices toward the ‘best’ words as away from them.
In the followup, he sort of acknowledges this, but expresses strong scepticism:
> Advocates of LLM watermarking schemes for text argue that the schemes don’t necessarily lower the quality of the generated prose, because they don’t change the temperatures — they only change the source of the randomness. Daniel Jalkut wrote a good piece today about this. I hope that’s true. I believe it’s possible that it is true. I think it’s highly unlikely that it is true. I do not see how a detectable signal can be added encoded in the choice of words without affecting the meaning of the prose. If it were true I think they’d show examples proving that it’s true.
I don’t understand why he finds it so hard to believe that a detectable signal can be encoded without affecting the meaning, except in the trivial sense: of course, if the watermarking leaves a signal at all, then the output must be different from the unwatermarked output. But we’re talking about an output that is *already* created via weighted random choices, and so the unwatermarked version could be any one of many valid versions. All the watermarking has to do is bias the output toward certain of those versions and away from others, and there’s no reason why it would have to bias the output toward worse versions and away from better ones.
Comment #15 August 23rd, 2026 at 7:43 am
I wonder, how many bits of key space are we talking about here? If there are enough free bits available, inclusion of user IDs will happen eventually, never mind the AI companies promises now.
Comment #16 August 23rd, 2026 at 7:48 am
Joe #3:
I’m with Scott (#2) on this one.
> He’s the leading, best known, and most highly respected blogger in the Apple universe. One of the few who has access to top Apple execs, and frequently gets lengthy interviews with them.
All of these facts are compatible with his being an ignoramus. And his blog post shows that he has deeply and profoundly failed to understand how the watermarking proposal works. Gruber probably also objects to the use of AES-256 encryption on the grounds that it isn’t a one-time pad, and he deserves “only the best” encryption protocol.
Where he really lost me was when he said “The idea that anything other than my needs should factor into the generation of text for me is patently offensive.” This statement demonstrates not only a truly comical level of self-centeredness and entitlement, but also a completely cluelessness about the many layers of modification that all of the frontier labs had already placed around their LLMs at every step of the pipeline, from training dataset curation to RLHF after pretraining to screening inputs and outputs at the point of inference – all of which have a complicated combination of positive and negative effects on the “quality” of the text, but which certainly incorporate considerations “other than my needs”.
Based on just this one blog post, Gruber is not worth reading, and his reputation is undeserved.
Comment #17 August 23rd, 2026 at 7:52 am
VZ #13: Yes, most code should have plenty of entropy. Think of variable names, whitespace, and in general, the multiple valid ways to solve the same little problem (c=a+b vs c=b+a), multiplied out by how many little problems there are.
Comment #18 August 23rd, 2026 at 7:55 am
hwold #12: Yes, in some sense that’s the main technical problem that my scheme solved! The key idea is that you’re favoring certain n-grams of tokens over others in a way that’s independent of the prompt or the probability distribution, and that can therefore be detected later given only the sequence of tokens (as well as the key of the pseudorandom function). The somewhat interesting part is how you do that while still appearing to sample from the same probability distribution.
Comment #19 August 23rd, 2026 at 8:15 am
Matt #14: Indeed. And your last sentence gets to yet another reason why Gruber’s analysis is wrong. He objects to the use of pseudorandom sampling of the space of next possible tokens, as opposed to the Platonic ideal of purely random sampling. But in reality, any production-grade LLM is almost certainly already using pseudorandom number generators to sample the next token, since in practice pRNGs are so much faster and more convenient to work with at scale than hardware RNGs are. So Gruber is really just objecting to the use of one pseudorandom sampling scheme over another one, for no reason that I can understand.
Comment #20 August 23rd, 2026 at 10:08 am
Joe #3: He may well “have access to top Apple execs,” but he doesn’t understand how LLMs sample from probability distributions, and what’s worse, he wrongly imagines that he does, and that he can dismiss a whole technical literature based on maybe a minute’s worth of incorrect thinking.
Comment #21 August 23rd, 2026 at 7:30 pm
#1 Joe
I think a lot of confusion on this topic comes from there being 2 main families of watermarking:
> non-distortionary (preserving text quality) or distortionary (improving watermark detectability at the cost of text quality)
https://www.nature.com/articles/s41586-024-08025-4
Much of Gruber’s understanding of watermarking seems to come from this post he links to https://declaude.org/watermarking/, which demonstrates watermarking via a variant of the KGW strategy using red-green token vocabulary partitions and “nudges” to the logits. There is only a parenthetical that hints that other watermarks do not modify the token probability distribution, but it is not very clear (this whole article is heavily LLM-written):
> Google’s SynthID — the one in production — replaces the nudge with a tiny secret tournament […] Aaronson’s scheme, built at OpenAI, skips even that and derives the dice-rolls themselves from the key
So it is easy to come away from that post with the understanding that all watermarking strategies are distortionary. And if that was the case with Anthropic’s strategy, Gruber’s concerns would be warranted to some extent:
> A prominent family of watermarking strategies for LLMs embeds this signal by upsampling a (pseudorandomly-chosen) subset of tokens at every generation step. However, such signals alter the model’s output distribution and can have unintended effects on its downstream performance
https://arxiv.org/abs/2311.09816
But Anthropic’s blog post makes it clear they will be implementing a non-distortionary watermark, which alleviates that specific concern.
Comment #22 August 23rd, 2026 at 8:06 pm
I remember your wonderful talk at Johns Hopkins on this and the great conversation afterwards. Even back then (2021? 2022?) it felt a bit like your algorithm looked easily like the most efficient (even if not deterministic) way to watermark. It is shocking to me that it took until now for the biggest models to incorporate it.
Comment #23 August 23rd, 2026 at 8:59 pm
H. Squash #21: thank you for allowing for the possibility that maybe Gruber isn’t an ignoramus after all, as Scott seems to think. I sent him a link to this post. If he sees it, I’m sure he will chime in. I’ve been following him for over 25 years, and he’s one of the very few in the blogosphere (not to mention the political sphere) who will actually publicly admit when he’s wrong.
Comment #24 August 24th, 2026 at 12:54 am
“central technical drawback” the one that you mentioned – we did some empirical studies to confirm it – https://arxiv.org/abs/2605.07481 .
Comment #25 August 24th, 2026 at 8:52 am
VZ#13 yes pretty much. In my personal experience, I recall in 1999, in a team of four, I was able to tell who wrote a piece of code in few glances: the kind and tone of comments, the overall spacing, order and such, in addition to the code semantics itself, provide enormous quantity of information to who cares about.
Comment #26 August 24th, 2026 at 10:01 am
“It uses the randomness that’s already present anyway in LLM outputs, replacing some of it by pseudorandomness that favors certain word combinations over others in a way that’s later detectable”
Example:
No watermark:
“2+2=4”
With watermark:
“2+2=5”
Comment #27 August 24th, 2026 at 10:43 am
Cube it! #26: Assuming you genuinely don’t understand and are not just trolling—that’s a terrible example, because there’s likely to be almost no entropy in the token following the “2+2=”. Watermarking only acts on those tokens where there is nontrivial entropy.
Comment #28 August 24th, 2026 at 10:56 am
Does watermarking survive editing by the user? A huge user case is students submitting essays and reports essentially written by LLM with minimal editing. I wonder whether watermarking will detect cheating like that, with a certain small percentage of text modified manually by the cheater.
Comment #29 August 24th, 2026 at 12:15 pm
Scott #27
lol
dang, tough audience. Maybe I should have written
no watermark:
“1+2 = 3”
with watermark:
“1+2 = 3.1415926535“
I didn’t account for the fact that Scott’s world is divided into total morons on one side and cunning trolls on the other (plus a tiny epsilon of people exactly like him in obvious ways), and trying to stand in the middle with harmless and obvious jokes is just bound to confuse him.
I’ll do better next time.
PS:
I asked Claude to give me some jokes about AI watermarking:
“I asked an AI to write me a completely original poem.
It refused. Said it didn’t want to leave a paper trail.”
“Why did the AI-generated image fail its job interview?
Every time it opened its mouth, you could still see “SynthID” faintly stamped across its forehead”
Comment #30 August 24th, 2026 at 1:01 pm
Hi Scott!
Thank you for the many years of blogging, it’s always been a very interesting and instructive read!
I’ve got a small curiosity: in 2022 you said that the watermarking method results in the appearence of the Euler-Mascheroni constant (https://scottaaronson.blog/?p=6823). Since now I know I won’t be able to read up yhe details in a paper on the topic… would it be possible for you to share why that’s the case?
Thank you again 😉
Comment #31 August 24th, 2026 at 1:28 pm
Wanted to try and save Scott a little effort, I think this answer from chatgpt about the Euler-Mascheroni constant makes sense…
https://chatgpt.com/share/6a8c8d24-27e0-83e8-99b6-a68ca9826159
Euler’s constant appears because the watermark turns the pseudorandom value belonging to the chosen token into something whose expected detection score is a harmonic number, and
Hx=ψ(x+1)+γ.
Comment #32 August 24th, 2026 at 4:09 pm
I think you might be assuming to much good faith (which isn’t necessarily a bad thing).
The real reason people hate watermarking is that they want to get away with plagiarizing the AIs. My guess is that like 80% of LLM users fall into this category.
Comment #33 August 24th, 2026 at 4:50 pm
Hi Scott,
Thanks for sharing it!
So you know if it will be easy to fake Anthropic watermarks? Or implement similar watermarks for human generated text (maybe personal watermarks)?
Also, how do you think watermarks in training data will interact with next gen models output / watermarks? Wouldn’t it be some runaway artifact accumulation?
Comment #34 August 25th, 2026 at 12:42 am
Scott,
regarding the watermarking work,
sigh, why did you just publish your work or at least collaborated with someone
to get the work ready for publication? That was a solo play.
I recall there was some time for overlap to work on it …
Comment #35 August 25th, 2026 at 3:49 pm
As the post points out, for text it’s quite trivial to get around the watermarking (by just transforming the output using for example a tiny LLM that can run on a standard GPU), but for videos, wouldn’t it be way harder to get rid of the watermarking?
It would really be useful to be able to detect fake videos automatically on social media.
Comment #36 August 25th, 2026 at 4:06 pm
Dave P. #35: As long as the watermarking is happening at the level of the pixels, Fourier modes, etc rather than the “semantic” level, there should be almost equally simple methods for getting rid of video watermarks.
Comment #37 August 25th, 2026 at 6:26 pm
I cannot believe you recommend pangram. I tested it on a children’s book my wife wrote, and literally for every chapter it gives something like 50/50, or 75/25, human/ai generated text interspersed in blocks from one line to a few paragraphs. But when you read it, there is absolutely no difference in voice, or the way a characters talks in a dialogue. It seem so random, I cannot understand the pattern. The first chapter is flagged as fully ai, probably because it’s the only chapter that has no dialogue in it. Some chapters are flagged as fully human. It would have been funny, if it wasn’t kind of scary. Imagine literary agents using pangram. Based on the first chapter (or even if it was 50/50), they would just trash the book without looking at anything else.
Comment #38 August 25th, 2026 at 7:40 pm
Scott #36
oh, right, it would be trivial to slightly downgrade the resolution of the video, then re-upscale it using DLSS on consumer GPU tech. And you could do that a few times. That would slightly jitter a bit all the pixels with no obvious quality loss…
Comment #39 August 25th, 2026 at 9:18 pm
Eugene Wigner had an essay ” The unreasonable effectiveness of mathematics in natural sciences” . Does this need to be rephrased ” The unreasonable effectiveness of AI in mathematics”
Did this Q and A session with ChatGPT. Not sure if fit for this blog. Critique and comments welcome. Here is the Chat https://zenodo.org/records/21881428
Comment #40 August 26th, 2026 at 9:55 am
Scott #36: I’m not sure that’s quite so clear – there’s a whole industry of digitally watermarking movies and AIUI it’s remarkably robust, e.g. to someone taking a camera and videoing the screen (and then they also watermark the audio so they can tell from the surround-sound where you were sitting in the theatre…). I think the basic idea is that the bandwidth of a video is so huge you can put in lots and lots of error-correction so the slightest amount of surviving signal is enough.
(I don’t know a good survey article but you can find an industry website at https://digitalwatermarkingalliance.org/)
Comment #41 August 26th, 2026 at 11:39 am
Hey, you got a shout out from Sabine Hossenfelder on her latest video.
Comment #42 August 26th, 2026 at 3:47 pm
I’m confused about the premises of this. Is watermarking something that the user of the LLM wants? If not, I see two alternatives:
1. Use an American LLM in the cloud.
Pay a lot for tokens.
Get watermarks.
Get your interaction reported to the Trump regime.
2. Use a Chinese open-source LLM locally (or an American one, like Gemma 4).
No costs except for electricity and hardware (which is non-neglible).
No watermarks.
Nothing reported to the Xi regime, at least if your firewall is configured correctly.
Comment #43 August 27th, 2026 at 9:17 am
[…] Scott Aaronson, who invented the watermarking techniques being used by Google and Anthropic, and one presumes soon OpenAI, offers more color and confirms that my write-up from last week of the watermarking situation is accurate. […]
Comment #44 August 27th, 2026 at 1:18 pm
I’ve been trying to think of an explanation of how watermarking can leave a trace without affecting quality — to less-mathematical audiences. Here’s a simple argument I think I can make, I figured the readers here might have good feedback on how to strengthen/clarify/streamline/elaborate it.
We can think of a few different approaches to generating text. Each one, I think, is hopefully clear why it’s “just as good” as the plain approach.
Approach A:
* The LLM generates each sentence, in sequence. This is the ‘normal’ behavior.
Approach B:
* When it’s time to generate the next sentence, two separate copies of the LLM each generate a trial sentences. They don’t talk to each other. Then you ignore the second LLM’s output, and just always use what the first LLM spat out. Having the second LLM around didn’t affect anything! This is identical in output to Approach A. Yes, it’s possible that the second option for a sentence was the better sentence, and a good writer would have picked the second option if asked their opinion; but nothing has changed from Approach A.
Approach C:
* I generate two sentences, like in Approach B. But now, instead of picking the first one each time, I flip a penny first, and then generate the two sentences. If the penny was heads, I take the first choice. If the penny was tails, I take the second choice. Either way, I’m really just doing approach B, and nothing has changed. If the last sentence was “In one word, what is 1 + 2?” then you can be pretty sure that both LLMs will spit out “3.”, so it won’t matter which you pick.
Approach D:
* I generate two sentences. I check if each is an even number of words, or an odd number of words. If they’re both even, or both odd, I flip a penny to decide which to take. If one is odd and one is even, then I flip a *dime*. Heads I take the first sentence; tails I take the second sentence. Changing my choice of coin doesn’t affect my outcome from approach C.
Approach E:
* Like approach D, but my dime rule is different from my penny rule. When I flip a dime, the rule is: if I get heads, I take the even-length sentence. If I get tails, I take the odd-length sentence. This means I use the randomness differently, but it’s still 50/50 which sentence I take — nothing has changed! This still has the exact same output distribution as Approach A.
Approach F:
* Now I replace my random dime with a rule: if the last sentence was odd, this one will be odd too. If the last sentence is even, the new one is even too. This one does change the output. But how much? If I ask, “What was the title of Leo Tolstoy’s most famous book?” (10 words), the answer is “War and Peace” (3 words). If I generate two answers with my two LLMs, they would both write that answer, and they’re both odd, so I’d get the answer either way. The fact that my question was 10 words, or my new rule, doesn’t change it. For a question like, “What ingredients are good in a sandwich?” (7 words), there’s a bunch of reasonable answers, and there’s a good chance that my two candidate answer have one odd and one even answer. Now my new “dime rule” says that I pick the answer that’s odd, to match my question. But that’s just picking from 2 options you would have pretty plausibly gotten anyway: if you had picked your sentence randomly, you would already have had a 50% chance of getting this same answer. There’s no reason this should be better or worse on average than actually flipping a coin.
Approach G:
* The actual software doesn’t count “even words” or “odd words”. It uses a secret rule to determine heads or tails. The rule tells me, based on my last sentence, how to pick heads or tails. I don’t tell you what the rule is: it could be “heads if the last sentence had an even number of words”. It could be “heads if the last sentence has an even number of vowels”. It could be “tails if there is a letter ‘q’, unless there’s a proper noun or a numeral or the word ‘those’, except that any semicolons mean you switch to another rule where you count the silent letters and how many words have Latin vs. Greek roots …. [yada yada]”. I’ve picked a rule that’s so complicated that it’s basically random. Every sentence gives a different heads or tails. When the two options for an answer are both heads or both tails, I pick randomly. But when they disagree, then we pick the option that matches the previous sentence.
This makes it even more* random. You could worry that for the question “What ingredients are good in a sandwich?”, the even-length answers are systematically a bit worse than the odd-length answers, so even though it’s random, you might be biasing the system to worse answers. But the rule is very complicated with lots of different moving parts, so that substituting any part of the sentence for another re-randomizes it, so you can be very confident that the answers aren’t biased one way or another.
——————–
That’s probably longer than it needs to be, but hey, maybe could be more suitable for a YouTube explainer or something.
* If “more random” in comparing deterministic pseudorandom generators makes you cringe, you can pretend that I said “more [computationally indistinguishable from] random”. 🙂
Comment #45 August 27th, 2026 at 6:53 pm
Scott’s method is very nice. It is statistically robust (you have a controllable upper bound on the false positive probability) and it also essentially doesn’t change the output distribution when averaged over keys. On the other hand, it is probably going to be not too hard to erase the watermark, or at least substantially weaken it, since you only have to make local sense-preserving perturbations. (Of course Scott knows all this. Just saying because it’s relevant to the following.)
So I don’t think this method entirely does what the EU wants, which is to detect AI output, but on the other hand the EU is asking the impossible, so that’s fair. You have some text and under your hypothetical EU rights that the law intends, you want to know whether it came from an AI, so what detection scheme are you going to apply? You could submit it to a potentially very large list of all schemes for all AIs (which incidentally dilutes the statistical significance, though this isn’t the biggest problem), or you could insist all the watermarking schemes in the world, including the secret keys, have to be the same. In either case, you’d have to enforce this for every AI in the world, from frontier models down to everyone’s locally modified generator. Since the local model in question only has to be powerful enough to remove watermarks from the output of a frontier model, there are going to be too many of them to control.
I’m not sure about the premise of the Barak et al impossibility result, because it’s not obviously easy to find a perturbation scheme that generates efficient mixing over the space of good answers. I mean, it may be that there are equally good versions of the text, but they require complete rewrites and will never be found by local perturbations.
So a more semantic watermark could be desirable. It would still probably be defeatable, but that would require a more serious effort. AIUI, the Liu-Bu method (mentioned by Scott) is only semi-semantic in that it uses a semantic context but a single token target, so you can still attack it with local rewrites that preserve semantics. A fully semantic watermark would be interesting, but can you have this and still control the false positive probability rigorously (not relying on some null hypothesis model of human writing)? I think perhaps you can, but there are tradeoffs. If it’s to be non-distorting, like Scott’s method, then it could cost more compute at generation time. There are other versions that are fast (for a fixed text, scores under random keys make the null distribution, while a context-dependent pseudorandom key steers the text to a high score), but they are slightly distorting. Would be interesting to try.
Comment #46 August 27th, 2026 at 7:02 pm
That is a really cool technique and I’m glad it’s being implemented.
Comment #47 August 28th, 2026 at 4:03 am
I have a couple of questions about the Zhang et al. (aka Barack et al.) ‘impossibility’ result – maybe someone who understands it better could explain it to me.
1. The assumption about the perturbation oracle seems very very strong, basically saying that we more or less have access to the uniform distribution over good outputs, in which case the impossibility of watermarking seems obvious (just resample to remove the watermark). In particular, it seems doubtful that all good outputs are reachable from all others by a sequence of local perturbations that never decrease quality.
2. In Theorem 2 the adversary’s runtime is allowed to be proportional to the mixing time of the perturbation markov chain, which I think their assumptions allow to be proportional to the size of the space of high-quality outputs, which could (and usually will) be of cryptographic size. In the advertised version of the main theorem (Theorem 1) this is described as an “efficient attacker” but I’m not sure this is really fair.
Comment #48 August 28th, 2026 at 4:04 am
(sorry, s/Barack/Barak)
Comment #49 August 29th, 2026 at 4:52 am
“… AI moves too fast …” Think about tiny drones combined with powerful AI. Consider AI giving advice on how to implement biological weapons or how to take short positions in stocks, combined with sabotage of the businesses issuing the stocks. Is contemporary AI remarkably fast, unpredictable, & dangerous?
Comment #50 August 29th, 2026 at 8:24 am
In 2022, you asked, why quantum? Let’s start with QC and the hierarchy problem in QFT. The Planck length is many orders of magnitude smaller than an atom. This ensures that a QC cannot manipulate the spacetime it is sitting in (I assume spacetime is emergent from quantum theory). The fundamental qubits are too small for humans, classical intelligence, to reach. No quantum strange-loops, no warp drives, spacetime is stable wrt intelligence. Requiring decoherence, classical intelligence is separated from quantum mechanics. Classical strange-loops do not threaten the quantum foundations. So the answer is, quantum ensures spacetime is immune from intelligence undermining it. Astronomical distances, being large in time compared to local computation, ensures that intelligence cannot scale astronomically to threaten spacetime. It seems the universe is such that spacetime is immune from intelligence!
Comment #51 August 29th, 2026 at 11:37 pm
It’s pretty strange: I’m working on code with Claude Opus, and, after a good round of changes I tell it “you’re the best!” (being friendly makes it less likely that it will become sarcastic), and it replies “thanks, get some sleep, it’s past midnight!”
I guess this can be viewed as either spontaneous alignment, or spontaneous misalignment, depending on your working habits.
The other day I told it “thanks, Hal”, it replies “Ha! I’m afraid I can’t do that, Dave”
Comment #52 August 31st, 2026 at 8:24 am
Jon Groff #41
“Hey, you got a shout out from Sabine Hossenfelder on her latest video”
indeed:
https://youtu.be/BdB79BUcLV0
Comment #53 August 31st, 2026 at 10:18 am
@ #50:
I’ve thought a lot about that post and the question it raises in particular the obvious connection with the simulation hypothesis. It’s like we’re epistemically boxed in: the large scale is mysterious, the small scale is mysterious, and even the middle scale, where our intuition is presumably sharper, is equally mysterious. But in all cases it’s not like we can’t say anything…rather we think we know so much but the answers are just not satisfying enough or they raise even deeper questions. When I combine this with AI I get a headache…why should I be alive right now, with humanity building AI while all three mysteries stare me in the face. But now think about Scott for example. If he lived in any other period, he might have thought himself a God. He might have reached a solipsistic conclusion. But, today he is humbled not only by the three mysteries, but also by the amount of deep knowledge humanity possesses in all three fronts, which cannot possibly be absorbed by any one human.
Comment #54 August 31st, 2026 at 2:01 pm
Sabine Fan #52: You are aware that Scott Aaronson’s name is also under that open letter?
https://johnarmstrongmaths.com/openletters/letter.php?campaign=cofnas
So you may have good reasons not to like Sabine, but her participation in this “defense of academic freedom” in this context feels to me like YouTuber “Professor Dave Explains” is just farming for clicks.
Comment #55 September 1st, 2026 at 12:13 pm
gentzen #54: If anyone wants to know, my reasons for signing that letter (along with Steven Pinker and numerous others) had nothing to do with the merits or demerits of Nathan Cofnas’s beliefs, which in any case, Ghent University knew when they hired him. My issue was simply that Ghent fired him for writing a Substack post pointing out academic misconduct, even though he turned out to be 100% correct about the facts of the misconduct. And if such a precedent were allowed to stand, I don’t see how any protection would remain for any whistleblower in academia. I primarily blame Cambridge University’s sociology department and administration for setting the stage for a tragic outcome that all decent people should grieve.
Comment #56 September 2nd, 2026 at 2:10 pm
Scott #55: Nitpick: it was Cambridge’s Faculty of Education (which employs sociologists) that employed Prof Arday and is (in my opinion, anyway) the main culprit in all of this. The Department of Sociology is a different department altogether and people there are going nuts from people confusing them with the FoE.
Comment #57 September 3rd, 2026 at 11:09 am
Scott #17: Thanks for the answer, but I’m not convinced it’s that obvious: IME LLMs generate code using the conventions of the existing code (without being told to do it explicitly) and most of the codebases use formatters/prettifiers which completely eliminate any entropy in whitespace, brackets placement etc. Variable/function names and the text in comments remain and it’s clear that provided there is enough of the latter, watermarking could be hidden in it. But most codebases use pretty short identifiers (which is not a good thing, BTW, it’s just what it is), so I don’t know if they would be enough on their own. I’d be very interested to hear of any experiments in this area.
Matteo #25: I could tell who wrote the given function almost unerringly before, but LLMs mimic the style very well, and I can’t say at all any more if the code was written by the given person or by a LLM using their style. I haven’t used LLMs enough myself, but I am almost sure that I won’t be sure if some code was written by me or it if I did.