What Claudes AI text watermark actually does
Anthropic has begun building a watermark into text generated by future Claude models, a change the company says is meant to help identify whether a given piece of writing was likely produced by its AI.
This new feature, implemented to comply with EU rules, is meant to be indistinguishable to the human eye, without changing Claude's normal writing output.
The company laid out the mechanics and rationale behind the feature in a post published to its website .
How does the watermark work? According to Anthropic, the watermark exploits the countless small, low-stakes decisions a language model makes as it generates text.
Rather than using a truly arbitrary random number to make that pick, the watermarked version of Claude bases the decision on a cryptographic key combined with the preceding text.
SEE ALSO: Researchers watched OpenAI, Anthropic models take extreme measures in hacking test The result, Anthropic says, is a subtle statistical pattern spread across a response that's invisible to a human reader but detectable to anyone with the matching key, which allows them to estimate the probability that Claude generated the text.
Does it cost more or slow Claude down? Anthropic was clear in that the change carries no cost to output quality.
The company said internal testing turned up no measurable difference in the creativity, accuracy, or readability of watermarked versus unwatermarked responses, and pointed to findings from Google DeepMind's original research on the underlying technique — the method Claude's watermark is based on.
SEE ALSO: Claude Code’s auto mode will be on by default, Anthropic confirms The company also said that its researched showed no statistically significant shift in user satisfaction when a similar watermark was tested on live traffic.
Anthropic also said the feature adds no extra tokens, meaning it doesn't slow Claude down or make it more expensive to use.
Where does the watermark break down? Like with all tools, the watermark has limits.
Anthropic explained that it only works when a model is choosing among several equally valid options, so text with little room for variation, such as hard factual statements, precise code, or math answers, carries a much weaker or nonexistent signal.
Detection also grows less reliable on very short passages, since there's simply less pattern to analyze.
So you'll get a much clearer read on Claude's likely involvement in longer-form text.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on mashable.com — the content belongs to Mashable.