Saturday, October 10, 2026 SourcesAbout🌓
🇺🇸 US ▾
BREAKING
› Dominion still has pending lawsuits against election deniers such as Rudy Giuliani and Sidney Powell› Russia is 'going backwards' in equipment and deploying post WWII-era tanks, according to Western officials› Podcast: One country musician is calling for other artists to oppose assault rifles› Bidets save you money and reduce waste — we tested the best options out there› 50+ products to make your life easier and our planet cleaner› Mother's Day is around the corner. Here are 50+ thoughtful gifts she'll love› A head-to-toe guide of how men should dress this spring, and where they should shop› 42 of the most useful travel products you can buy on Amazon› The 7 best high-yield savings accounts of April 2023› Taxes are due tomorrow. Here's how to file for an extension› Dominion still has pending lawsuits against election deniers such as Rudy Giuliani and Sidney Powell› Russia is 'going backwards' in equipment and deploying post WWII-era tanks, according to Western officials› Podcast: One country musician is calling for other artists to oppose assault rifles› Bidets save you money and reduce waste — we tested the best options out there› 50+ products to make your life easier and our planet cleaner› Mother's Day is around the corner. Here are 50+ thoughtful gifts she'll love› A head-to-toe guide of how men should dress this spring, and where they should shop› 42 of the most useful travel products you can buy on Amazon› The 7 best high-yield savings accounts of April 2023› Taxes are due tomorrow. Here's how to file for an extension
Technology

LLMs respond differently to harmful prompts when AI watermarking is used

Ars Technica ·
LLMs respond differently to harmful prompts when AI watermarking is used

In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate.

Anthropic recently disclosed its future Claude models will use SynthID-Text , an approach Google created and released as open source.

It uses a secret key that subtly changes the process a model uses for choosing the next word in a sentence.

Whereas a top next word choice might be “cloudy,” the key might change it to “overcast.” Anyone who knows the key can determine if it was generated by the platform using it.

New research shows that SynthID-Text can change not just word selection but also the tools a model invokes and the chances it will adhere to or disregard safety guardrails it has been trained to follow.

The threat can become greater in the face of an adversarial prompt, in which an attacker attempts to cause a model to carry out a harmful action, such as revealing a password or other sensitive information.

Instructions that normally wouldn’t be followed will, in some cases, be performed once the watermarking is deployed.

The finding underscores the need for developers to thoroughly test how their LLMs and agents behave when watermarking is in place.

Changing safety behavior “As compared to the same models without watermarking, it is definitely going to change their behavior, especially when we place it under adversarial conditions, or we make these models call tools when they’re powering an agent,” Andrea Siposova, an AI security researcher at Lasso Security, told Ars.

“Watermarking is made to not be perceptible to a reader, but we know that when we are changing anything about what the model is generating, it is going to cause some tradeoffs, it’s going to show up somewhere.” Read full article Comments

Read the full article on Ars Technica ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on arstechnica.com — the content belongs to Ars Technica.

More from Ars Technica

See all ›

More in Technology

See all ›