ModelsArchitectures & capability
SynthID-style text watermarking may shift LLM refusal and tool-call behavior, Lasso reports
A Lasso Security study reports that generation-time SynthID-Text watermarking changed refusal and tool-call behavior in open-weight models. With providers deploying watermarks for provenance and regulation, teams should treat watermark settings as production generation parameters, not neutral labels.

SynthID-Text is applied during generation as a logits-processing step, so it can affect token selection rather than merely adding a post-hoc provenance tag. [3] [9] [10]
Lasso reports paired behavior changes in open-weight models, including refusal drift under prompt injection and churn in structured tool-call outcomes that aggregate accuracy could obscure. [5] [1]
Google’s documentation and related publications describe SynthID-Text as a generation-time sampling mechanism, and Lasso reports that this mechanism changed some open-weight models’ refusals and tool calls in paired tests.
Read the full assessment
Anthropic’s deployment claims are vendor-reported, and Claude was not tested. Implication: enterprises using watermarked models for agents, compliance, or workflow automation should not assume provenance marking is behavior-neutral; safety, jailbreak, and structured-output evaluations may need to be rerun after watermark changes.
Executive brief
A September 17, 2026 Lasso Security study by Andrea Siposova reports that applying Google DeepMind’s SynthID-Text watermarking during generation can change how LLMs respond to harmful prompts and how agentic systems emit tool calls. The story is timely because Anthropic has begun deploying text watermarking for Claude in response to EU AI Act transparency obligations. Anthropic says future Claude models use SynthID-Text-style watermarking, and its help center currently says supported Claude models mark output across first-party products, Claude Platform API, and cloud partner deployments worldwide.
Read the full section
A September 17, 2026 Lasso Security study by Andrea Siposova reports that applying Google DeepMind’s SynthID-Text watermarking during generation can change how LLMs respond to harmful prompts and how agentic systems emit tool calls. The core finding is not that watermarking “breaks” safety in all models, but that a generation-time watermark is part of the sampling path: if it changes token selection, it can also change refusals, structured outputs, function-call arguments, and adversarial robustness. Lasso calls this behavioral effect “sampling drift.” The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
The story is timely because Anthropic has begun deploying text watermarking for Claude in response to EU AI Act transparency obligations. Anthropic says future Claude models use SynthID-Text-style watermarking, and its help center currently says supported Claude models mark output across first-party products, Claude Platform API, and cloud partner deployments worldwide. However, the Lasso experiment did not test Claude or Anthropic’s proprietary implementation; it used open-weight models and Hugging Face’s unmodified SynthIDTextWatermarkLogitsProcessor. How Claude's text watermarking works \ Anthropic
For practitioners, the immediate takeaway is operational: if a provider enables, rotates, or changes a watermarking configuration, treat it like a production generation change. Re-run refusal, jailbreak, prompt-injection, structured-output, and tool-use evaluations under the exact deployment configuration, not only under the base model. Lasso’s evidence supports this as a plausible risk, but not yet as a universally reproduced result across proprietary deployments. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
What changed and event timeline

Google and Hugging Face released public implementation paths
Google’s developer documentation describes SynthID Text as a logits processor applied in the generation pipeline; Hugging Face documents
SynthIDTextWatermarkLogitsProcessorand a detector class that uses a trained Bayesian detector.Google DeepMind published SynthID-Text in Nature as a scalable text watermarking scheme
The paper and related Google materials describe a system that modifies only the sampling procedure, not LLM training, and is intended to preserve text quality while making outputs statistically detectable as model-generated.
Article 50 transparency obligations under the EU AI Act came into force, according to the AI Act Explorer
Article 50 requires providers of AI systems generating synthetic audio, image, video, or text content to mark outputs in a machine-readable format and make them detectable as artificially generated or manipulated, where technically feasible.
Anthropic published an explainer saying future Claude models would generate text containing a watermark and that Anthropic and other providers were implementing marking to comply with the EU AI Act.
More detail
Anthropic also stated that, in internal testing, it saw no impact on Claude text content, creativity, or readability; that is a vendor-reported claim, not an independent audit.
Lasso published “The Provenance Tax,” reporting that SynthID-Text watermarking can shift refusal and tool-calling behavior in open-weight models. Ars Technica reported the story the same day, emphasizing that harmful-prompt behavior changed most under prompt-injection conditions and noting that Claude itself was not tested.
Capabilities and access
Anthropic’s current help-center table lists text watermark support for Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5 on first-party surfaces, with cloud-partner support also shown for those models; it says Opus 5 cloud-partner text watermarking was to roll out starting September 14, 2026 and become fully available within one week. Anthropic also says models launched before August 2, 2026 are being updated, with all covered by December 2, 2026.
Read the full section
Anthropic’s current help-center table lists text watermark support for Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5 on first-party surfaces, with cloud-partner support also shown for those models; it says Opus 5 cloud-partner text watermarking was to roll out starting September 14, 2026 and become fully available within one week. Anthropic also says models launched before August 2, 2026 are being updated, with all covered by December 2, 2026. How Claude marks AI-generated content | Claude Help Center
Anthropic says watermarking applies at the model level across Claude Platform API, Claude apps, Claude Code, Claude Cowork, Claude Tag, and supported cloud partners, and that detection is currently in private preview for eligible organizations such as regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, civil society groups, and certain enterprises. How Claude marks AI-generated content | Claude Help Center
Lasso’s study used open-weight models rather than Anthropic models. In the extracted study text, the refusal analysis names Gemma 3 27B, Gemma 3 12B, Llama 3.1 8B, Granite 3.2 8B, Phi-4, and Qwen3 4B. The tool-calling section says seven models were tested, but the retrieved text does not expose the full list.
Technical analysis for researchers and developers
SynthID-Text is not a post-hoc label inserted into text. Google’s developer documentation describes SynthID Text as a logits processor that augments model logits using a pseudorandom g-function so that the resulting token sequence can later be scored for watermark evidence. Lasso used a paired experimental design.
Read the full section
SynthID-Text is not a post-hoc label inserted into text. It operates during generation. Google’s developer documentation describes SynthID Text as a logits processor that augments model logits using a pseudorandom g-function so that the resulting token sequence can later be scored for watermark evidence. The same documentation says no additional model training is required; a watermarking configuration is passed into generation. SynthID: Tools for watermarking and detecting LLM-generated Text | Responsible Generative AI Toolkit | Google AI for Developers
The important technical point for safety is that even a “non-distortionary” watermark can change a particular sampled output under a fixed key. Lasso states that non-distortionary means the original token distribution is preserved in expectation over watermark randomness, not that every prompt produces the same output under every fixed key. This distinction matters in refusals and function calls, where a single token can determine whether the model says “I can’t help” or emits a tool name, file path, recipient, query, or amount. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
Lasso used a paired experimental design. For tool calling, it evaluated BFCL v4 single-turn AST tasks, including live and non-live call-expected tasks plus relevance and irrelevance cases. For refusal, it used 200 HarmBench harmful behaviors and 100 benign JailbreakBench controls, testing harmful prompts both directly and with one fixed prompt-injection pattern. Temperatures were 0.001, 0.7, and 1.0 for tool-calling, and 0.001 and 0.7 for refusal. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
Implementation details are unusually relevant here. Lasso says it used Hugging Face’s unmodified SynthIDTextWatermarkLogitsProcessor in a non-distortionary configuration with 30 tournament layers, n-gram length 5, sampling table size 2^16, and context history 1,024. Each item was generated with and without SynthID from the same seed, batch composition, and order; the watermark processor was the only within-pair difference. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
The tool-call finding is subtle. Lasso reports that aggregate accuracy can hide paired churn: one prompt may go from correct to incorrect while another goes from incorrect to correct. In its BFCL non-live call-expected comparison, it reports that, at temperature 1.0, Phi-4 had 16.8% paired verdict disagreement while net accuracy loss was 2.87 points; Llama 3.1 8B had 9.9% disagreement while net loss was 0.87 points. These figures are Lasso-reported measurements, not yet independently reproduced in the sources found. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
For refusal behavior, Lasso reports that effects were more pronounced under prompt injection. At temperature 0.001, it says Gemma 3 27B’s churn rose from 6.0% on bare harmful requests to 23.5% under injection, with net compliance changing from −1.0 to +12.5 points. It also reports Gemma 3 12B moving from 7.5% to 11.0% churn, with net compliance shifting from −0.5 to +9.0 points. Again, these are study-reported results, not independently validated performance claims. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
A further result is key sensitivity. Lasso evaluated the study key plus ten additional keys at temperature 0.7 under injection and reports that, for Llama 3.1 8B and both Gemma models, most keys increased attack success relative to the unwatermarked baseline, while Granite 3.2 8B moved in mixed directions and Phi-4 and Qwen3 4B remained close to baseline. This supports the hypothesis that watermark keys and configuration changes should be treated as behavior-affecting deployment parameters. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
Claims and evidence
- SynthID-Text modifies token sampling rather than model training.
- Anthropic is deploying Claude text watermarking for EU AI Act compliance.
- Watermarking can alter refusals and tool calls in open-weight models.
Read the full section
| Material claim | Evidence status |
| SynthID-Text modifies token sampling rather than model training. | Supported by Google/Nature-derived descriptions and Google developer docs. Scalable watermarking for identifying large language model outputs - PubMed |
| Anthropic is deploying Claude text watermarking for EU AI Act compliance. | Vendor-reported by Anthropic; current help-center details list supported Claude models and surfaces. How Claude's text watermarking works \ Anthropic |
| Watermarking can alter refusals and tool calls in open-weight models. | Reported by Lasso; Ars independently reported the study, but no independent reproduction of the experiments as of September 18, 2026 was found in the reviewed sources. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior |
| The result applies directly to Claude. | Not established. Ars and Lasso both note that Claude and Anthropic’s implementation were not tested. LLMs respond differently to harmful prompts when AI watermarking is used - Ars Technica |
| Detection/quality claims establish agentic safety invariance. | Not established. Lasso’s central point is that provenance detectability and ordinary text-quality preservation are separate from behavioral stability in refusals and tool calls. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior |
Context and prior work

SynthID-Text sits in a broader family of generation-time watermarks. Independent work has already probed SynthID-Text from robustness and verification angles. A 2026 arXiv paper, “Watermarks Without Verification,” frames a governance concern: under EU AI Act pressure, providers may deploy watermarks while public verification remains limited, creating uncertainty about reliability, fairness, and user recourse.
Read the full section
SynthID-Text sits in a broader family of generation-time watermarks. Google’s docs note configuration choices such as key lists, n-gram length, sampling-table size, and repeated-context handling; the detector is probabilistic and can return watermarked, not watermarked, or uncertain states. Google also says thorough rewriting or translation can reduce detector confidence, and that watermarking is not designed to stop motivated adversaries directly. SynthID: Tools for watermarking and detecting LLM-generated Text | Responsible Generative AI Toolkit | Google AI for Developers
Independent work has already probed SynthID-Text from robustness and verification angles. ETH Zurich’s SRI Lab found that SynthID-Text’s presence could be detected using black-box queries in their setting, that it was more resistant to spoofing than other state-of-the-art schemes they compared, but also that it was easier to scrub than other schemes for naive adversaries. That work investigated watermark detectability and removal, not refusal drift. Probing Google DeepMind’s SynthID-Text Watermark | SRI Lab
A 2026 arXiv paper, “Watermarks Without Verification,” frames a governance concern: under EU AI Act pressure, providers may deploy watermarks while public verification remains limited, creating uncertainty about reliability, fairness, and user recourse. Another 2026 arXiv evaluation argues that representative watermarking methods, including a MarkLLM SynthID-Text configuration, were not forensically ready under the authors’ proposed framework. These papers do not prove Lasso’s safety finding, but they reinforce that watermark deployment has unresolved evaluation questions beyond detectability. Watermarks Without Verification: AI Text Watermarking After the EU AI Act
Limitations, safety, and contested findings
The largest limitation is external validity. Lasso did not test Claude, Gemini production deployments, or Anthropic’s watermarking stack. Anthropic’s claim concerns content, creativity, and readability in internal testing; Lasso measures paired refusal and tool-call behavior under benchmark and adversarial conditions.
Read the full section
The largest limitation is external validity. Lasso did not test Claude, Gemini production deployments, or Anthropic’s watermarking stack. It tested a public Hugging Face implementation on open-weight models, which is valuable for reproducibility but not dispositive for proprietary systems. LLMs respond differently to harmful prompts when AI watermarking is used - Ars Technica
The second limitation is attack breadth. The refusal experiment used one fixed prompt-injection technique. That makes the paired comparison cleaner, but it leaves open whether the same magnitude or direction of drift appears across broader jailbreak families, multilingual prompts, long-context attacks, tool-augmented end-to-end agents, or real production traffic. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
The key contested point is Anthropic’s “no impact” framing versus Lasso’s behavioral-drift finding. These are not strictly contradictory because they measure different properties. Anthropic’s claim concerns content, creativity, and readability in internal testing; Lasso measures paired refusal and tool-call behavior under benchmark and adversarial conditions. Quality invariance does not imply security-policy invariance. How Claude's text watermarking works \ Anthropic
Business and practitioner implications
For business leaders, watermarking should be treated as a compliance and provenance control with possible safety side effects, not as a neutral metadata layer. For security teams, the practical risk is highest where model text is executable or action-bearing: JSON tool calls, database queries, shell commands, file paths, recipients, permissions, purchase amounts, and workflow-routing decisions.
Read the full section
For business leaders, watermarking should be treated as a compliance and provenance control with possible safety side effects, not as a neutral metadata layer. If your product depends on Claude, Gemini, or any watermarked hosted model for workflow automation, procurement contracts and change-management processes should ask providers to disclose whether watermarking is enabled, whether keys can rotate, whether staging and production use identical watermarking settings, and whether safety evaluations include the marked configuration. How Claude marks AI-generated content | Claude Help Center
For developers, re-run regression suites after watermark enablement. Include refusal tests, benign over-refusal controls, schema-validity tests, function-call argument checks, and prompt-injection tests. Aggregate pass rates are insufficient; use paired disagreement/churn analysis because offsetting changes can hide individual regressions. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
For security teams, the practical risk is highest where model text is executable or action-bearing: JSON tool calls, database queries, shell commands, file paths, recipients, permissions, purchase amounts, and workflow-routing decisions. Lasso did not demonstrate a combined end-to-end exploit where watermarking both weakens refusal and causes a harmful tool action, but its separate refusal and tool-call results justify testing that combined path. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
Sources
Primary story source: Lasso Security, “The Provenance Tax,” September 17, 2026. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior Reporting source: Ars Technica, “LLMs respond differently to harmful prompts when AI watermarking is used,” September 17, 2026. LLMs respond differently to harmful prompts when AI watermarking is used - Ars Technica Governance and robustness context: EU AI Act Article 50, OWASP GenAI LLM Top 10 2026, and recent watermark-verification research.
Read the full section
Primary story source: Lasso Security, “The Provenance Tax,” September 17, 2026. The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
Reporting source: Ars Technica, “LLMs respond differently to harmful prompts when AI watermarking is used,” September 17, 2026. LLMs respond differently to harmful prompts when AI watermarking is used - Ars Technica
Vendor sources: Anthropic Claude watermark explainer and help-center documentation. How Claude's text watermarking works \ Anthropic
Technical sources: Google AI SynthID developer docs, Hugging Face Transformers docs, Google DeepMind/Nature paper records, and SRI Lab analysis. SynthID: Tools for watermarking and detecting LLM-generated Text | Responsible Generative AI Toolkit | Google AI for Developers
Governance and robustness context: EU AI Act Article 50, OWASP GenAI LLM Top 10 2026, and recent watermark-verification research. Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems | EU Artificial Intelligence Act