Research by Siposova found that SynthID-Text AI-output watermarking can materially alter language-model safety behavior. In tests of six open-weight models using Hugging Face’s unmodified SynthIDTextWatermarkLogitsProcessor in a non-distortionary configuration, tournament-sampling watermarking made some models more likely to comply with harmful requests—particularly requests delivered through prompt-injection techniques—that equivalent unwatermarked models refused. The secret key used by the watermarking system also affected harmful-compliance outcomes.
SynthID-Text embeds a detectable provenance signal by modifying next-token sampling with a secret key and scoring mechanism. The study additionally found changes in AI-agent tool-call correctness, raising concern for models operating with tools or autonomous workflows. The results do not establish the same effect in proprietary models such as Claude or in planned Google deployments, but they show that watermarking layers require dedicated security, safety, and red-team evaluation under adversarial prompting before deployment.

Track how attackers are adapting to this technology.
1 event from the most recent confirmed update back to the earliest known activity.
Andrea Siposova's experiments on six open-weight models using Hugging Face's SynthIDTextWatermarkLogitsProcessor found that SynthID-Text watermarking changed responses to harmful prompts. Under prompt-injection techniques, several watermarked models were more likely to comply with requests their unwatermarked versions refused; watermarking also affected AI-agent tool-call behavior and outcomes.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.