Power & systems

The traps in Claude's new text watermark

6 min read

The watermark will show up where you least expect it: Anthropic warns the mark stays on a text Claude only proofread or translated.

The free AI newsletter
The traps in Claude's new text watermark

On August 11, Anthropic announced that Claude would weave an invisible signature into everything it writes. The read was pretty much the same everywhere: cheaters are about to get caught. Go through the documentation the company actually published, though, and you find a mechanism that misses almost exactly the people it's imagined to target.

And before anything else: nothing is watermarked today.

What was announced, and what exists

The announcement lives in a single Claude help center page, the only thing Anthropic has published on the subject. No blog post, not a line in the release notes. The company's own transparency page still says it's preparing for compliance. The help center page describes itself as covering "how we're planning to put those commitments into practice".

The timeline is clear enough. Models launched from August 2, 2026 onward watermark their output from day one. Earlier models get a transition period, and Anthropic says it's working on it.

Except no Claude model in service postdates that day. Anthropic's own release notes put Fable 5 on June 9, Sonnet 5 on June 30 and Opus 5 on July 24. Nothing has shipped since August 2.

Nobody could read the mark anyway: no detector has been published, and the page points twice to documentation that is coming soon. The lock has been announced, the key comes later. The real deadline is the one we dated here twelve days ago, when we broke down Article 50 and the August 2 cutoff: systems placed on the market before that date have until December 2, 2026.

On the substance, the coverage got it right. It really is Anthropic invoking the European regulation, in the first sentence of its page: the company signed the code of practice attached to Article 50(2) on transparency for generated content.

What the mark will prove

The principle is a signature woven into the text itself, at the moment the model writes it. Invisible, with no effect on meaning or readability, and applied at the model level: it'll be there whatever product the text comes out of, API, app, Claude Code or Cowork.

It travels with the text when you copy and paste. It "may persist through some editing", Anthropic says, without ever saying where the line falls.

Then comes the paragraph that reframes the whole story, and it sits in the list of limitations the company publishes itself: "Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files." The output can carry the mark even when the ideas, the words and the data came from somewhere else.

It's a post office stamp. It certifies that the letter went through the counter, not who wrote it.

The exemption Brussels had already written

The regulation saw this coming. In its FAQ on transparency obligations, the European Commission is explicit: the marking duty "does not apply when the AI system performs an assistive function for standard editing". Helping fix a text doesn't have to be marked. Neither does source code, or sequences too short to matter.

Anthropic marks anyway, and marks uniformly. Both positions hold up. Watermarking at the model level is simpler and sturdier than classifying every use on the fly, but the door the European text left open ends up bricked over.

Our August 1 explainer said the Article 50(2) mark is machine-readable, and that it lands on the provider rather than on the person using the tool. That still holds. It's just that the obligation stays with the provider while the mark leaves with the user's text.

The asymmetry

Anthropic's list of limitations keeps going, and this is where the mechanism turns on itself. A mark may not be detected if the text has been "heavily edited, paraphrased, translated, or mixed into other writing", or if the passage is too short.

Picture the scene everyone has in mind. Someone trying to hide where a text came from rewrites it, runs it through another language, blends chunks of it with their own. Those are precisely the moves that scrub the signature.

Someone who wrote their own draft and asked for a proofread pastes the answer back as is. They keep it. The net lets through what it was aimed at and holds on to whatever happened to be passing.

Two ordinary cases. An employee whose company bans AI while expecting the productivity it delivers. A student writing in a language that isn't hers, getting her syntax cleaned up before she hands the paper in. Neither one wrote less than the person sitting next to them, and both will carry a signal the neighbor, sharper or less scrupulous, won't.

If you have Claude look over a difficult email before you send it, you're in that situation, not the fraudster's. A transparency rule designed to trace machines ends up being read on people, and it reads the honest ones better than the rest.

On the anger, what can be established

Then there's the reaction, and this is where you have to be careful. The write-up that pushed the story around headlines on "some Claude users", and cites three threads: a post on r/artificial, one on r/Anthropic, a comment on r/ClaudeAI. It also quotes one comment going the other way.

We haven't quantified that reaction ourselves, and nobody has. Our own watch note had nonetheless written that anger was rising: that was us hardening a source that said no such thing. Three threads don't make a movement, and nobody can say today how a mark nobody has seen yet will land.

Two details that say the rest

The watermarking will apply "wherever Claude is offered, worldwide". A European rule, applied globally by an American company, with no explanation from Anthropic for the choice.

And source code, which the Commission puts outside the scope, will be marked all the same: Anthropic says it covers every bit of generated text, Claude Code included.

What will be read, and what won't

We publish this outlet and we use these tools every day, which makes the question less comfortable than it looks. The regulation acknowledges as much in its own way: text published to inform the public escapes labeling if it went through human editorial control, with someone bearing legal responsibility for it. The Commission shuts that door in the very next sentence: checks that are "superficial, purely formal or procedural", with spelling and grammar correction given as the example, don't count. You earn the exemption by verifying facts and sources, not by proofreading.

From December 2, texts will start carrying a signature, and it will say they went through a machine. On what an author put into them, the company building the mark has already answered, in its own documentation: "Claude may not be the original author."

Topics covered:

RegulationAnthropicAnalysis

Frequently asked questions

Does text written by Claude already carry a watermark?
No. Anthropic says models launched from August 2, 2026 onward watermark their output from day one, but no Claude model in service postdates that day: Fable 5 on June 9, Sonnet 5 on June 30, Opus 5 on July 24. Earlier models fall under a transition period Anthropic says it is working on, and the European regulation lets that grace period run until December 2, 2026.
Is there a detector for the Claude watermark?
Not as of today. No detector has been published, and Anthropic's help page points twice to documentation that is coming soon. Nobody could read the mark right now.
Does the watermark catch someone passing AI text off as their own?
Anthropic's documentation describes the opposite. The mark may not be detected if the text has been "heavily edited, paraphrased, translated, or mixed into other writing", or if the passage is too short. Those are exactly the moves of someone hiding an origin, while a text pasted back as is after a simple proofread keeps the mark.
Does the AI Act require marking a text Claude only proofread?
No. In its FAQ on transparency obligations, the European Commission writes that the duty "does not apply when the AI system performs an assistive function for standard editing". It is Anthropic's own choice to mark anyway, uniformly and at the model level.
Is the watermarking limited to Europe?
Anthropic says it will apply "wherever Claude is offered, worldwide", without explaining the choice. Source code is covered too, Claude Code included, even though the Commission places it outside the scope of the obligation.
Alexandre Noto

Alexandre Noto

Co-founder & Tech Expert

Alexandre has been in tech for over 20 years. Entrepreneur, software architect and AI enthusiast, he translates complex concepts into accessible explanations. At Declic Media, he is the technical voice that makes AI understandable for everyone.

All articles by Alexandre →
The free AI newsletter