Anthropic says future Claude models will carry a Claude text watermark embedded in how the model selects words, rather than in visible labels, hidden characters or added text metadata. The company announced the approach on August 14, saying it is intended to help meet the European Union’s AI Act transparency requirements.
The mark is designed to indicate the likelihood that Claude participated in writing a passage. That is a narrower claim than proving who authored it, a distinction that will matter wherever a positive detection is treated as a verdict rather than evidence.
How does Claude’s text watermark work?
Language models produce text one token at a time, estimating possible next words from the preceding context. For a primer on that process, see how LLMs generate an answer. Anthropic says its watermark works when more than one next-word choice would be appropriate and would preserve the sentence’s meaning.
In Anthropic’s example, a sentence ending in “The weather today was cold and…” may plausibly continue with “overcast” or “grey.” The ordinary selection involves randomness. With watermarking, Anthropic says, the system instead uses a key and several preceding words to determine that otherwise random choice.
Repeated across a longer response, those low-stakes selections create a statistical pattern. Readers will not see it in any individual word. Someone with Anthropic’s key can test whether a passage’s sequence of choices matches the keyed process, then assign a probability that Claude was involved.
Anthropic describes its system as a version of SynthID-Text, the approach Google DeepMind published in a 2024 Nature paper. The company has not said it is using an identical implementation, so claims about its exact technical configuration would be guesswork.
What can a Claude watermark show?
According to Anthropic, verification can address only whether text was likely written in part by Claude. It cannot establish that a passage was human-written, identify writing from a different AI system, or show that Claude originated the underlying ideas or all of the prose.
The evidence also weakens where the model has little room to choose. Short samples provide fewer word decisions, so Anthropic says detection performs poorly on them. Confidence rises with longer passages. Factual writing can yield a sparser mark because accuracy may leave only one sensible wording. The same problem applies to light proofreading: the watermark covers words Claude chooses, and a lightly edited human draft may contain too few model-made changes to register.
What Anthropic says about quality and cost
Anthropic says watermarking does not force Claude to select words the model would not otherwise consider. It says internal tests found no effect on content, creativity or readability, while citing Google DeepMind’s SynthID-Text research as reporting no statistically significant quality difference in its tests.
Separately, Anthropic says the method adds no tokens, costs users no more and has no practical effect on output quality or content. Those are company claims, not independent findings in the announcement.
Anthropic says the EU has required AI providers serving its market to mark AI-generated content since August 2, and that several major providers are making similar changes. Its announcement, however, speaks of future Claude models and does not establish the full rollout scope.
This story draws on original reporting from The Verge.