Anthropic says text watermarking scheme relies on inconsequential words
In an effort to "watermark" text that Claude has generated and comply with the EU AI Act, Anthropic unveiled a plan on Friday to modify its bots' choice of words in a way that would be detectable as the product of an AI.
Traditional watermarks are patterns or images overlaid on currency, postage, or official documents as an assertion of authenticity.
In the digital realm, the term is more flexible and can refer to a variety of techniques for applying an identifier to electronic data.
Anthropic's approach involves influencing inconsequential word choices made by its models, a technique introduced in Google DeepMind's SynthID-Text paper.
To oversimplify things, large language models are fancy autocomplete engines which work by predicting the next word in a sequence of words.
Anthropic explains that while composing sentence output like "The weather today was cold and…" a model like Claude might respond with words like "cold" or "gray" and would be unlikely to respond with a word like "sugary." That's the theory, but when actually asked to complete that sentence, Claude Opus 4.8 went a bit overboard: "…crisp, the kind of cold that nips at your fingertips and turns your breath to little clouds.
The sky was a pale, washed-out blue, and everything felt sharp and clear." And then it checked to see if users thought that was useful, asking, "Want me to take it somewhere specific — cozy, gloomy, cheerful? Or keep going with the same tone?" But remove whatever training has been applied to promote engagement and simulate literary style, and that's basically what Claude is doing here – predicting the next word in a sequence.
Anthropic asserts that in most cases, the example sentence could be completed by either "cold" or "gray" and "the meaning of the sentence is largely the same either way." The watermark is generated by deviating from the predicted word to something else.
A different source of randomness is used and that can be detected with a digital key.
As Google DeepMind researchers explain in their paper: "Generative watermarking works by carefully modifying the next-token sampling procedure to inject subtle, context-specific modifications into the generated text distribution.
Such modifications introduce a statistical signature into the generated text; during the watermark detection phase, the signature can be measured to determine whether the text was indeed generated by the watermarked LLM." Anthropic insists this will be done with low-stakes passages in a way that won't alter the meaning.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.