Saturday, 10 October 2026 SourcesAbout🌓
🇬🇧 UK ▾
BREAKING
Technology

AI model watermarking changes agent behavior

The Register ·
AI model watermarking changes agent behavior

Watermarks that European law requires be added to AI-generated content to establish provenance may come at a cost.

According to Lasso Security, AI model watermarking changes how AI agents handle tools and safety refusals.

The altered behavior isn't necessarily worse but can be, particularly under adversarial prompt injection.

With the implementation of the EU AI Act, providers of AI models must mark the output of their software with machine-readable code.

Google DeepMind's SynthID-Text is one method for doing so, and has been adopted by Anthropic and by OpenAI.

The benefit of this sort of digital labeling is that manipulative or deceptive AI-generated content can be more easily detected, even if it does have the potential to stigmatize the usage of AI.

Anthropic's explanation of how it applies watermarks to Claude output involves intervening in the prediction that results in specific words.

For example, if Claude were emitting the sentence "The weather today was cold and…" then it might favor one statistically likely candidate (e.g.

"overcast") over an alternative (e.g "gray").

It may be possible to detect those additions.

"Watermarking is designed for provenance, but SynthID-Text changes the process by which the model generates each next token," Lasso explained in a blog post provided to The Register.

"At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection." "Watermarking uses low-stakes choices like these – which occur many times over a piece of generated text – to leave a pattern in Claude’s responses," Lasso Security added.

"That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it." While a reader might not notice the word choice bias, AI agents can be subtly sensitive to vocabulary differences.

Lasso found that this sort of digital content tagging can affect tool calling and refusal behavior.

Read the full article on The Register ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.

More from The Register

See all ›

More in Technology

See all ›