Anthropic will watermark text from future Claude models with a version of Google DeepMind’s SynthID-Text, using keyed word choices rather than hidden characters or extra tokens. The 2024 study introducing SynthID-Text compared feedback on nearly 20 million watermarked and unwatermarked Gemini responses and found no statistically significant difference in user ratings, though the test covered Google’s implementation rather than Anthropic’s planned configuration. Its signal depends on the amount and kind of text Claude produces, limiting what a detection result can establish.
How It Works
- Claude will encode a statistical pattern through keyed word choices, without hidden characters or extra tokens.
- A 2024 Gemini study covering nearly 20 million responses found no statistically significant difference in user ratings.
- Short, factual or lightly edited passages may contain too little signal for reliable detection.
- Independent SynthID-Text tests found paraphrasing weakened detection, but they did not test Claude's implementation.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
How the key shapes word choice
A language model produces text one word or token at a time. At each step, it assigns probabilities to possible continuations based on the preceding text. Some decisions have only one sound answer. Others leave room for several words that preserve the meaning.
After the words The weather today was cold and, either overcast or grey could make sense, while sugary would not. The watermark operates only within that set of acceptable, low-stakes choices.
SynthID-Text operates in that second category. A secret key, combined with a small amount of preceding text, supplies the randomness used to settle among acceptable candidates. The choices remain within the model’s existing range. Once they accumulate, someone holding the key can test whether the sequence is consistent with the keyed process.
The technique changes sampling rather than model training.
Detection depends on available choices
The signal grows when Claude produces longer, less constrained prose because each eligible decision adds evidence. Short passages offer fewer observations. Factual sentences, equations and working code may give the model little freedom to choose a different token without creating an error.
Light proofreading has a similar problem. If Claude changes only punctuation and a handful of words, most of the returned passage was never selected by the model. The resulting mark may be too sparse for detection.
A detector would evaluate how closely the text matches choices associated with Anthropic’s key and return a likelihood of Claude involvement. A found mark means Claude may have processed the material. It cannot separate original generation from heavy editing, and it cannot establish human or AI authorship. Failure to find a mark also does not establish human origin.
Heavy editing, paraphrasing or later translation can disrupt the sequence. Mixing marked text with other writing can dilute it.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
Google’s system yielded to paraphrasing
Tests from ETH Zurich’s SRI Lab, built on its 2024 watermark-stealing research, give the main counterexample. Google’s production deployment was unsuitable for thousands of similar queries, so the researchers used a local model watermarked with SynthID-Text. They tested off-the-shelf paraphrasers and reported extremely high scrubbing success even without first stealing the watermark.
The result shows how meaning can remain while the token sequence changes enough to break detection. It does not show that Claude’s watermark has been defeated. The public disclosure does not say whether Anthropic will use the same settings, scoring methods or safeguards within the broad SynthID-Text approach.
Separate experiments found paraphrasing, copy-and-paste changes and back-translation weakened detection, but those tests did not measure Claude’s implementation.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
The public test is still missing
No public Claude detector or technical specification is available as of Aug. 15, 2026, so outsiders cannot test the implementation directly. Anthropic plans to offer a detection API, with technical details still forthcoming.
Anthropic intends the API to let third parties check for a Claude watermark. The disclosure does not specify a minimum passage length, measured survival rates under editing or an accepted detector error rate.
Frequently Asked Questions
What does Claude's text watermark add?
Nothing visible. It changes the sampling process used to choose among plausible next words, leaving a statistical pattern without extra tokens or hidden characters.
Does a detected watermark prove Claude wrote the text?
No. It indicates Claude may have processed the text and cannot distinguish original generation from heavy editing or establish human authorship.
Why are short or factual passages harder to detect?
They offer fewer acceptable word choices, giving the key fewer opportunities to create a measurable pattern. Code and equations can be similarly constrained.
Can editing remove the mark?
Heavy editing, paraphrasing, translation or mixing with other text can weaken detection. Independent studies did not test Anthropic's implementation.
When can users check a Claude watermark?
Anthropic says it plans a detection API for third parties, but has not published a release date, minimum passage length or accepted detector error rate.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
Related stories


IMPLICATOR