In 2024, Nikola Jovanović, Robin Staab and Martin Vechev of ETH Zurich's SRI Lab spent less than $50 querying a public API to reverse-engineer the rules behind a text watermark. They then scrubbed the pattern from generated text.
Anthropic will apply Europe's AI-content marking requirement to Claude worldwide.
Anthropic will embed watermarks in text generated by Claude models launched after the rule took effect in Europe and attach signed provenance data to supported files, under a policy described in an Aug. 10 support document. The policy covers the Claude app, its API, Claude Code, Claude Cowork and Claude Tag, including supported models reached through AWS, Google Cloud and Microsoft Foundry.
What Changed
- Anthropic will embed watermarks in text from Claude models launched after Aug. 2, 2026, and attach C2PA provenance metadata to supported files, applying the policy wherever Claude is offered rather than only in the EU.
- Article 50(2) of the EU AI Act became applicable Aug. 2, 2026, with fines of up to 15 million euros or 3 percent of worldwide annual turnover. Systems already on the market may get until Dec. 2, 2026.
- Published research has repeatedly removed watermarks from other systems: paraphrasing tools stripped Google's SynthID-Text in more than 90 percent of attempts, and UnMarker cut SynthID image detection from 100 percent to about 21 percent.
- Anthropic has published no technical specification and no detection tool, so outside researchers have had nothing to test.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
What the code requires
Article 50(2) of the EU AI Act became applicable on Aug. 2, 2026. It requires providers to make synthetic text, images, audio and video machine-readable and detectable as AI-generated or manipulated. Noncompliance after the rule took effect can bring fines of up to €15 million or 3 percent of worldwide annual turnover. A provisional agreement reached in May 2026 would give systems already on the market until Dec. 2, 2026, to comply with the marking requirement.
Anthropic signed Section 1 of the voluntary Code of Practice, which maps that duty into specific commitments for providers. By the end of July 2026, Section 1 had 82 signatories. The full code had about 190 participating organizations.
The code prescribes no single technical answer. No available method meets its requirements for effectiveness, interoperability, robustness and reliability on its own. Legal researcher Natalia Garina, who holds an LL.M. in Digital Law from the Catholic University of Lyon, examined those four requirements in a June 2026 analysis of the final code. The European Commission and AI Board have assessed the code as adequate, and two task forces are due to start work in September 2026.
How the marks work
Published text-watermarking schemes generally create a statistical pattern as a model writes. The software slightly biases which words the model picks, leaving a pattern across a long enough passage that a detector can measure even though a reader cannot see it. Anthropic has not disclosed whether Claude's method works this way. The company says its pattern will travel when text is copied and pasted and may survive some editing.
Supported files such as SVG, PNG and JPG images will receive signed metadata conforming to the C2PA provenance standard. The label records that Claude processed the file and can show whether someone later tampered with it. Open-source software already exists to remove C2PA metadata. Format conversion, re-saving and screenshots can also strip it.
Anthropic has published no technical specification or detection tool, so outside researchers have had nothing to test. The company says detection details are forthcoming and cautions that a detected mark does not prove Claude created the underlying material. Heavy editing, paraphrase or translation may erase a text signal; a very short passage may never contain enough evidence.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
What researchers removed
The research below does not measure Anthropic's system. The ETH work tested Google's text watermark, and the Waterloo work tested image marks rather than Claude's text method or its signed C2PA metadata.
Jovanović, Staab and Vechev built the black-box watermark-stealing attack for an ICML 2024 paper. With a budget of under $50 in API costs, the team reverse-engineered watermark rules through public API queries. Spoofing a state-of-the-art scheme then succeeded at over 80 percent, and scrubbing climbed from 1 percent to over 80 percent against a best prior baseline below 25 percent.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Their team's later tests of Google's deployed SynthID-Text found that ordinary paraphrasing tools removed the mark in more than 90 percent of attempts at a false-positive rate of one in 1,000. After black-box queries first exposed the pattern, removal approached 100 percent. Forging the mark was harder, succeeding in 4 to 15 percent of attempts.
Andre Kassis, a computer science PhD candidate at the University of Waterloo in Ontario, and Urs Hengartner, an associate professor there, built UnMarker to attack image watermarks. Their paper appeared at the 46th IEEE Symposium on Security and Privacy in May 2025. Across seven image-marking schemes, the best detector still recognized only 43 percent of processed images. In a later test, UnMarker cut Google's SynthID image detection from 100 percent to about 21 percent.
The choice to ship
By early August 2024, OpenAI had kept a text-watermarking tool for ChatGPT undeployed for more than a year because opinion inside the company was divided. A survey commissioned by OpenAI found support for tools that identify AI content outnumbered opposition four to one. A separate survey found 30 percent of ChatGPT users would use the service less if it carried watermarked content.
Anthropic chose deployment with stated limits. A mark may survive copying. Revision can make it disappear. Anthropic says a detected mark indicates content may have been processed by Claude and is not fully conclusive. Failure to find one does not establish that AI played no part.
The two EU task forces scheduled for September 2026 will begin work with providers already committed to marking output and researchers already documenting how marks fail. "We always rush to develop these tools and our excitement overshadows the security aspects," Kassis said. "We only think about it in hindsight and that's why we're always surprised when we find out how malicious attackers can actually misuse these systems."
Frequently Asked Questions
What exactly will Claude mark?
Text generated by supported Claude models will carry an embedded watermark. Supported files such as SVG, PNG and JPG images will receive signed provenance metadata conforming to the C2PA standard.
Does this apply only to users in the EU?
No. The policy covers the Claude app, its API, Claude Code, Claude Cowork and Claude Tag, including supported models reached through AWS, Google Cloud and Microsoft Foundry, wherever Claude is offered.
What does the EU rule actually require?
Article 50(2) of the AI Act, applicable since Aug. 2, 2026, requires providers to make synthetic text, images, audio and video machine-readable and detectable as AI-generated or manipulated. Noncompliance can bring fines of up to 15 million euros or 3 percent of worldwide annual turnover.
Can these watermarks be removed?
Research on other systems says yes. Paraphrasing tools removed Google's SynthID-Text in more than 90 percent of attempts, and open-source software already exists to strip C2PA metadata. None of that research tested Anthropic's method, which has not been published.
Does a detected mark prove Claude wrote something?
No. Anthropic says a detected mark indicates content may have been processed by Claude and is not fully conclusive. Failure to find a mark does not establish that AI played no part.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR