Anthropic has announced that future Claude models will include a statistical watermark designed to help determine when the assistant was involved in creating or modifying a text.

The company says the system will not identify the person who used Claude or reveal which organization or conversation produced the content. Its purpose will be to estimate the likelihood that Claude was involved.

How will it work?

Language models create text by selecting words or parts of words from several possible alternatives. The watermark will influence some of those choices to leave a statistical pattern that readers will not be able to see.

A detector with access to the corresponding key will be able to analyze the pattern and estimate whether Claude participated in the content. Anthropic says its system will be based on a version of SynthID-Text, a technology originally developed by Google DeepMind.

It will not be absolute proof

The signal may be harder to detect in short texts, highly factual passages, code, or documents where Claude only made minor corrections.

Anthropic also acknowledges that a complete rewrite can remove the signal. A detection result should therefore not be interpreted as definitive proof of authorship, fraud, or plagiarism. It would only indicate that Claude was probably involved at some point.

Why is Anthropic implementing it?

The announcement is connected to transparency requirements under the European Union’s Artificial Intelligence Act. These include measures for marking and facilitating the detection of certain content generated or manipulated using artificial intelligence.

Anthropic says it will apply the watermark globally when it launches the corresponding models. The company also plans to provide an API for checking the signal, but it has not yet announced when it will become available or what its access conditions will be.