ADAM PULSE Knowledge Base
AI · Content Provenance · Governance · EU AI Act

AI text watermarking, and the difference between AI-detected and AI-authored

Short answer

As of August 2026, text generated by supported Claude models carries an invisible statistical watermark, and Anthropic has published how it works. It is not a hidden character, it is not metadata, and it is not a list of "AI words." It is a pattern in which words the model picked, spread across the whole passage.

Three things matter for a business, and all three are commonly got wrong:

The regulatory driver is the EU AI Act, Article 50, whose transparency obligations became applicable on 2 August 2026. The marking applies worldwide, not only in the EU.

What actually changed, and when

The rule

Article 50 of the EU AI Act requires providers of generative AI systems to mark their outputs so that machines can recognise them:

Outputs of AI systems (audio, image, video, text) are marked in a machine-readable format and detectable as artificially generated or manipulated.

The obligations became applicable from 2 August 2026. To help providers show compliance, the European Commission finalised a Code of Practice on Transparency of AI-generated Content on 10 June 2026 and subsequently confirmed it as an adequate voluntary tool for demonstrating compliance. By late July 2026 roughly 190 organisations had signed it.

The Code requires that technical solutions be "effective, interoperable, robust, and reliable as far as technically feasible" — wording that matters later in this article, because it is an engineering-limits standard, not a guarantee.

One transition is still running, and it is narrower than it is usually reported. The Article 50(2) marking and detection requirement gives applicable systems that were already placed on the market before 2 August 2026 until 2 December 2026 to comply. That transition covers those systems and that obligation only — the rest of Article 50 is already in force.

The obligations carry weight. Breaches of Article 50 attract administrative fines of up to €15 million or 3 percent of total worldwide annual turnover for the preceding financial year, with proportionality considerations for SMEs and small mid-cap companies.

Anthropic's implementation

Anthropic's support documentation states the scope plainly:

Claude models launched on or after August 2, 2026 support marking at launch.

and, for everything released before that:

We're also working to add marking support to Claude models released before that date.

On geography, there is no EU-only carve-out:

Marking will apply to output from supported models wherever Claude is offered, worldwide.

Two distinct mechanisms are in play, and conflating them causes most of the confusion:

The two marking mechanisms Anthropic describes, and why they behave differently.
What is generated How it is marked What that means in practice
Ordinary generated text An imperceptible watermark woven into the text itself Travels with copy and paste; survives some editing; nothing to strip out of a file
Supported generated files (.svg, .png, .jpg) Signed provenance metadata using the C2PA open standard A cryptographic credential attached to the file; a different technology with different failure modes

Marked is not the same as required to be marked

This is the part most coverage misses, and it matters more than the mechanism.

Article 50(2) does not require marking for everything a model touches. Outputs are exempt where the AI performs an assistive function for standard editing or does not substantially alter the input data or its meaning — the Commission's own worked example is grammar correction. Machine-to-machine exchanges never seen by a person, and short sequences of numbers, symbols or source code, are also outside it.

Anthropic marks anyway. Its own documentation is explicit that people use Claude to "proofread, translate, summarize, or convert files" and that the output "can carry a Claude mark even if the underlying ideas, text, or data originated from another source."

So there are two different sets, and they are not the same set:

What Anthropic marks, set against what Article 50(2) actually requires to be marked.
  Marked by Anthropic Required to be marked by Article 50(2)
Claude drafts the document Yes Yes
Claude translates your writing Yes Yes — every word of the output is the model's
Claude fixes your grammar Yes, where enough signal survives No — assistive standard editing
Claude reformats your text Yes Generally no — meaning is not substantially altered

A model-level watermark cannot tell wholesale generation from a comma fix, because at the moment of generation there is nothing to tell it with. The consequence is the one this whole article turns on: a detected mark cannot be read as an authorship finding, and it cannot even be read as evidence that the law required the content to be marked at all.

The part that was unknown for about a fortnight

When this topic first started circulating, the honest position was that Anthropic had confirmed that it watermarked text but had not disclosed how. That is no longer the case. On 14 August 2026 Anthropic published a technical explanation, and it is specific enough to reason about.

This matters for anyone repeating earlier commentary: a good deal of writing published in the first half of August is now out of date, including the widely repeated hedge that the mechanism was undisclosed.

How words become a watermark

What a language model is doing when it writes

A language model does not compose a paragraph internally and then type it out. It produces one token at a time — a token being a word, part of a word, or a piece of punctuation — and at each step it holds a set of candidates with probabilities attached.

Take a sentence with an obvious gap:

The clouds became increasingly ______.

dark, grey, heavy, ominous, dense are all defensible. The model has to pick one, and ordinarily it picks using a source of randomness.

The trick: change the randomness, not the choice

That last word is the whole mechanism. Anthropic's description is that the watermark alters the source of the randomness rather than the range of acceptable answers. Instead of an arbitrary random draw, the system uses a secret key together with a few preceding words to settle which of the equally good candidates gets picked.

Anthropic is explicit about the constraint it works under:

[the system does not] push Claude to choose a word it wouldn't have considered anyway.

So the output stays natural. No individual word is evidence of anything — grey is just grey. But run that same key-driven selection across hundreds of choices and the sequence stops looking arbitrary. Anyone holding the key can test whether the words are consistent with the choices that key would produce. Anyone without it sees ordinary prose.

That is why a watermark detector and an "AI detector" are answering completely different questions:

Style inference and watermark detection are frequently confused, and are not the same evidence.
  What it asks What it needs What a positive result means
AI classifier Does this text resemble AI writing? Nothing but the text The style matched a model of AI style
Watermark detector Does this text carry the specific signal from a known key? The key The text is statistically consistent with that key's choices

Where the idea came from

Anthropic describes its method as a version of the SynthID-Text approach published by Google DeepMind, in a paper that appeared in Nature in 2024. The broader family traces back to a 2022 proposal by Scott Aaronson, and shares one design principle throughout: the watermark only changes the source of the randomness used to pick among words.

This lineage is worth knowing for one practical reason. The approach is described in a peer-reviewed paper rather than existing solely as a vendor claim, so it can be studied, tested and argued with. That is unusual and it is a point in its favour.

It is also the correction worth making loudly: this is not a rumour attributed to "recent reporting." Anthropic itself names the SynthID-Text lineage in its own technical write-up.

What the watermark is not

It is not a hidden character

People hear "invisible watermark" and picture zero-width Unicode wedged between the words. That technique exists in general, and stripping it is trivial — but it is not what is happening here. The signal is carried by which words were chosen, so there is nothing sitting between them to delete.

It is not metadata

For ordinary generated text there is no field, header, or attached record. The C2PA provenance metadata Anthropic describes applies to supported generated files, which is a separate mechanism. Converting text to plain text removes formatting, HTML and document metadata; it does not touch the word sequence.

There is no list of "AI words"

delve, tapestry, crucial, leverage — none of these is a watermark. A human can write any of them; a model can avoid all of them. Vocabulary-spotting is folklore, and it produces false accusations against people who simply write formally.

The em dash is not a watermark

It is punctuation. It has been in use for centuries. The association between em dashes and AI output is style-based inference, and treating it as proof of provenance is a category error — one that has already been used to accuse people wrongly.

What removes it, and what does not

This is the section most readers actually came for, so it is worth being precise rather than reassuring.

What is known about watermark survivability, per Anthropic's own documentation.
Action Effect on the text watermark
Copy and paste Survives — Anthropic states the mark travels with copied text
Paste into Notepad / strip to plain text No effect on the word sequence, so no reason to expect removal
Change the font, size, colour, spacing No effect — presentation does not rewrite language
Export to PDF Text signal unchanged; a file may additionally carry C2PA metadata
Fix a typo, change a comma, edit a sentence Likely survives — may survive some editing
Substantial rewriting Progressively degrades
Complete rewrite, every word replaced Removed — Anthropic states this plainly
Heavy paraphrasing Weakens, potentially to the point of non-detection

There are also cases where the signal is thin or absent from the outset, which is different from being removed:

Those last two deserve emphasis together, because commentary has generally got them backwards. The widespread worry has been "I only asked it to proofread, and now I will be falsely accused." Anthropic's account points the other way: a light proofread may well be undetectable. Meanwhile a translation — which many people would describe as their own work in another language — is fully marked.

What a detection actually proves

Anthropic's own framing is the ceiling on any claim anyone else makes:

A watermark can only determine that Claude was likely involved with the content at some point.

And, on the limit of that inference, the mark cannot distinguish "Claude wrote this" from "Claude heavily edited this." It does not confirm human authorship. It does not identify whether some other AI was involved. And it carries no identifying information traceable to a user or an organisation.

That last point deserves its own line for business readers: the watermark does not embed your account, your company or your client. If you have been worried that AI-assisted text carries a tag identifying who generated it, it does not.

So a detection supports one narrow claim — this text is statistically consistent with having passed through Claude — and no more. It does not answer:

There is no public detector yet

Anthropic has said it will offer a watermark detection API and is working out the implementation details.

Until that exists, treat any product claiming to detect Claude's official watermark as claiming something it cannot do, because detection requires the key. What such a service is actually running is a style classifier — the same technology as before, with better marketing. If your organisation is about to buy AI-detection tooling, ask directly: does this test for a watermark using a provider key, or does it classify writing style? The answers carry completely different evidentiary weight.

The mirror image is worth the same scepticism. Tools advertising that they remove AI watermarks are selling against a signal they cannot measure — they have no detector either, so they cannot show you a before and after. Per Anthropic's own account the only reliable removal is a complete rewrite in which every word is replaced, at which point the question of what the tool contributed rather answers itself.

Why this matters to a business

Watermarking is usually discussed as a schools-and-essays story. The governance problem lands hardest on organisations, because businesses now run AI through marketing copy, proposals, contracts, research, customer service, documentation, code, email and knowledge bases — often without a policy that defines any of it.

The unhelpful question is "was AI used?" The useful questions are narrower:

The obligation that lands on you, not on Anthropic

Everything above is a duty on the provider of the model. Article 50(4) puts a separate duty on the deployer — the organisation publishing the output. That is the half that applies to your business, and it is narrower than people assume.

The rule: AI-generated or manipulated text published to inform the public on matters of public interest must be clearly labelled. Not marketing copy. Not internal documents. Not a novel. Content published to inform the public on a matter of public interest.

And there is an exception with a condition attached. The labelling duty does not apply where the text has undergone human review or editorial control and a natural or legal person holds editorial responsibility for the publication.

The condition is where organisations will get caught. The Commission's guidance is that such review "must be substantive and not limited to superficial matters or cursory approval." A spell check does not clear it. A skim and an approval click does not clear it.

That gives readers a distinction worth writing into a policy:

A human looked at it and a human exercised substantive editorial judgment over it are not the same claim — and only the second one carries the exception.

Note the shape of this. On the provider side the instinct is to mark everything. On the deployer side the label is required only in a narrow case, and a named accountable human removes it. A wholly AI-drafted article on a matter of public interest can carry an invisible watermark and correctly carry no reader-facing label, because an editor took responsibility for it. Those two facts are not in conflict; they are answering different questions.

Do not write a policy that equates detection with authorship

The trap is a policy that says AI-detected means AI-authored. That sentence is now demonstrably false, and a policy built on it will produce unfair outcomes in both directions: it will accuse people who used a translator, and clear people who had a model draft the whole thing and then rewrote it.

A more useful policy distinguishes:

And it sets disclosure by consequence rather than by tool:

Disclosure calibrated to consequence, rather than a blanket “AI was used” declaration.
Content type Reasonable standard
Internal email and notes No disclosure needed
Customer-facing technical documents Named human reviewer before release
Contracts, legal or regulatory content Human approval, recorded
Public communications and marketing Editorial accountability, as with any published work

For anyone handling an accusation

If someone is accused on the basis of a detection result, the questions to put to whoever is making the accusation are:

  1. What does your tool actually test for — a provider watermark using a key, or writing style?
  2. What is its false-positive rate, and on what kind of writing was that measured?
  3. How long was the sample? Short passages carry little signal.
  4. Was the passage largely factual? The mark is sparser there.
  5. What does the vendor itself claim the result proves?

A tool that cannot answer the first question is not evidence of provenance. It is evidence that the writing resembled a model of AI writing, which is a much weaker thing, and which describes a good deal of competent formal prose written by people.

Does every AI system do this now?

Not identically, and the differences matter.

Google developed SynthID at DeepMind, covering images, video, audio and text, with the text method described in the Nature paper Anthropic cites. That is the most publicly documented approach in the field.

Anthropic applies the text watermark described above to supported Claude models, plus C2PA metadata on supported generated files.

For any other provider, the honest answer is: check that provider's own current documentation. Whether a given system marks its output is a product decision, and it should never be inferred from the fact that a competitor does. Roughly 190 organisations signed the EU Code of Practice, but signing a code and having shipped a particular technical mechanism are different facts, and a signature is not a detector.

The direction of travel is clear enough. The industry question is shifting from can we guess whether AI wrote this? to can the generating system provide machine-readable evidence about its own involvement? The second question is far more tractable, and it is the one regulation is now pointing at.

A word on C2PA and files

For images and other supported file types, the mechanism is not statistical at all. C2PA — the Coalition for Content Provenance and Authenticity standard — attaches signed provenance metadata to the file. It is a cryptographic credential rather than a pattern in the content.

The trade-offs run opposite to the text watermark:

The two approaches fail in different directions, which is largely why both exist.
  Text watermark C2PA file metadata
Where the signal lives In the word sequence In a signed metadata record
Removed by Rewriting the language Stripping metadata, re-encoding, screenshotting
Survives copy and paste Yes Not as plain content
Tells you Statistical likelihood of involvement A verifiable, signed assertion

In practice a screenshot of an image destroys C2PA provenance, while a screenshot of text destroys nothing about the text watermark — retyping or OCRing the same words preserves the word sequence. Neither mechanism is a serial number etched into metal, and neither vendor claims it is.

What "written by AI" is going to mean

The uncomfortable conclusion of all this is that the categories are wrong, not the technology.

Traditional authorship assumes a clean chain: a person thinks, writes, edits and publishes. A realistic 2026 document looks more like this:

Ask "who wrote it?" and the question does not have an answer, because it was built on an assumption that no longer holds. A watermark detection on that document would be correct and nearly useless: yes, Claude was involved at some point.

This is why organisations that write a single binary rule — AI-generated content is not permitted — find it unenforceable within a quarter. The rule cannot survive contact with the ordinary use of a spellchecker that happens to be a language model.

The better framing is accountability rather than provenance. Someone's name is on the document. That person is answerable for whether it is accurate, whether it is theirs to publish, and whether the process behind it complied with whatever rules apply. That standard has worked for every previous authoring tool and it does not need replacing now.

What we would do about it

For most businesses this changes very little today and quite a lot within a year. The proportionate response:

Two things remain genuinely unknown, and we would rather say so than fill the gap. Anthropic has not published a timeline for the detection API, and it has not published false-positive or false-negative rates for the watermark. Asked about both directly by Ars Technica in August 2026, the company responded with a statement that did not address either question. Until those numbers exist, nobody — including us — can tell you how often a detection would be wrong.

Frequently asked questions

Does Claude put a watermark in the text it generates?

Yes. Anthropic states that Claude models launched on or after August 2, 2026 support marking at launch, and that it is working to add marking to models released before that date. The marking applies to output from supported models wherever Claude is offered, worldwide.

Is the AI watermark a hidden character?

No. It is not zero-width Unicode or any inserted character. Anthropic describes a watermark woven into the generated text itself, carried by which words the model selected rather than by anything placed between them.

Is the watermark stored in metadata?

Not for ordinary generated text. Anthropic separately applies signed C2PA provenance metadata to supported generated files such as .svg, .png and .jpg. Those are two different mechanisms with different properties.

How does a text watermark actually work?

Anthropic describes it as altering the source of the randomness the model uses to choose between equally good word options. A secret key plus the preceding few words determines which candidate is selected. Anyone holding the key can test whether a passage is consistent with the choices that key would produce.

Is Claude's watermark based on Google's SynthID?

Anthropic describes its method as a version of the SynthID-Text approach published by Google DeepMind in a 2024 Nature paper. The wider family of techniques traces back to a 2022 proposal by Scott Aaronson.

Does copying and pasting remove the watermark?

No. Anthropic states that the mark travels with the text when it is copied and pasted, and may survive some editing.

Does pasting the text into Notepad remove it?

No. Converting to plain text strips formatting, HTML and document metadata, none of which carry the signal. The word sequence is unchanged.

Does changing the font remove an AI watermark?

No. Font, size, colour, bold, italic and paragraph spacing are presentation. They do not rewrite the language that carries the signal.

Does saving as a PDF remove it?

Not from the text. If the words are unchanged the statistical signal remains. Files may separately carry C2PA provenance metadata, which is a different mechanism.

Can editing remove an AI watermark?

Light editing may not. Anthropic states that a complete rewrite in which every word is replaced will remove it, and that substantial rewriting degrades it progressively.

Does paraphrasing remove an AI watermark?

Heavy paraphrasing can weaken or eliminate it, because the sequence of words carrying the signal is being replaced. This is a fundamental limitation of statistical text watermarking rather than a flaw in any one implementation.

Is AI-generated code watermarked?

No, where exact output is required. Anthropic states the watermark is not applied in those cases. Changing a token in code could break compilation, alter behaviour or introduce a vulnerability, so correctness takes priority over detectability.

If Claude only proofreads my writing, will it be detected?

Possibly not. Anthropic states that proofreading changes might not be enough to make Claude's involvement detectable. This is the opposite of the common worry, and it matters for anyone drafting an AI policy.

If Claude translates my writing, is the translation watermarked?

Yes. Anthropic states that watermarking does apply to translation, because every word of the output is chosen by Claude.

Does a short AI answer carry a watermark?

Barely. Anthropic states detection does not work well on small samples, where there are fewer word choices and therefore less information to test.

Is the watermark weaker in factual writing?

Yes. Anthropic states that watermarking is sparser on factual passages, because there are fewer word choices available that do not reduce the accuracy of the text.

Does a detected watermark prove Claude wrote the content?

No. Anthropic states a watermark can only determine that Claude was likely involved with the content at some point, and that it cannot distinguish Claude wrote this from Claude heavily edited this.

Does the watermark identify who generated the text?

No. Anthropic states the mark carries no identifying information traceable to users or organisations. It does not embed your account, your company or your client.

Can I detect Claude's watermark today?

Not yet. Detection requires the key. Anthropic has said it will offer a watermark detection API and is working out the implementation details. Until then, any service claiming to detect Claude's official watermark is running an ordinary style classifier instead.

Should I buy a tool that removes AI watermarks?

There is no way for such a tool to demonstrate that it worked, because no public detector exists for it to test against. Anthropic's own account is that reliable removal requires a complete rewrite in which every word is replaced.

Is an em dash an AI watermark?

No. It is punctuation that predates computing by centuries. Associating it with AI output is style-based inference, not provenance detection.

Are certain words AI watermarks?

No. There is no list of AI words. A human can write any word a model favours, and a model can avoid all of them. Vocabulary-spotting produces false accusations against people who write formally.

What is the difference between an AI detector and a watermark detector?

An AI detector asks whether text resembles AI writing and needs nothing but the text. A watermark detector asks whether the text carries a specific signal from a known key, and needs that key. They carry very different evidentiary weight.

Why are AI companies watermarking text now?

Article 50 of the EU AI Act requires machine-readable marking of AI-generated output, and became applicable on 2 August 2026. The European Commission finalised a Code of Practice on Transparency of AI-generated Content on 10 June 2026, which around 190 organisations had signed by late July 2026.

Does the watermark only apply in Europe?

No. Although the driver is EU regulation, Anthropic states that marking applies to output from supported models wherever Claude is offered, worldwide.

What should our company AI policy say?

Distinguish AI-authored, AI-assisted, AI-edited, AI-translated and AI-researched rather than asking only whether AI was used, and set disclosure by consequence: none for internal notes, a named reviewer for customer-facing documents, recorded human approval for contractual or regulatory content.

Someone was accused of AI use based on a detector. What should we ask?

Ask what the tool tests for, a provider watermark using a key or writing style; its false-positive rate and on what writing that was measured; how long the sample was; whether the passage was largely factual; and what the vendor itself claims a positive result proves.

Does my company have to label AI-generated text that it publishes?

Only in a narrow case. Article 50(4) of the EU AI Act requires a clear label on AI-generated or manipulated text published to inform the public on matters of public interest. It does not apply where the text has undergone substantive human review or editorial control and a natural or legal person holds editorial responsibility. The Commission is explicit that the review must be substantive and not limited to superficial matters or cursory approval, so a spell check does not qualify.

Does a Claude mark mean the law required that content to be marked?

No. Article 50(2) exempts outputs where the AI performs an assistive function for standard editing or does not substantially alter the input or its meaning. Anthropic marks that content anyway, because a model-level watermark cannot distinguish wholesale generation from a light edit at the moment of generation. Marked and legally required to be marked are different sets.

What happens on 2 December 2026?

That is the end of the transition for the Article 50(2) marking and detection requirement as it applies to systems already placed on the market before 2 August 2026. It applies to those systems and that obligation only. The rest of Article 50 has been in force since 2 August 2026.

What are the penalties for getting Article 50 wrong?

Administrative fines of up to 15 million euro or 3 percent of total worldwide annual turnover for the preceding financial year, with proportionality considerations for SMEs and small mid-cap companies.

How often does the watermark produce a wrong answer?

Nobody has published that. Anthropic has not released false-positive or false-negative rates, and did not answer when asked directly by Ars Technica in August 2026. Any vendor quoting you an accuracy figure for Claude watermark detection today is quoting something they cannot have measured.

References

Anthropic's own documentation and the European Commission throughout. Every statement in this article about what the watermark does is Anthropic describing its own system, not a third party characterising it — a distinction that matters here more than usual, because a good deal of early commentary was written before the mechanism was published.

  1. Anthropic — How Claude marks AI-generated content— source for the August 2, 2026 model cutoff, the worldwide scope, the C2PA metadata on supported .svg, .png and .jpg files, the statement that a detected mark means content may have been processed by Claude, and the note that detection details are forthcoming.
  2. Anthropic — How Claude's text watermarking works— published 14 August 2026. Source for the mechanism (altering the source of randomness in word selection using a key and the preceding words), the SynthID-Text lineage, the Scott Aaronson 2022 antecedent, the planned detection API, and the stated limits on short samples, factual passages, code, proofreading and translation.
  3. European Commission — Code of Practice on Transparency of AI-generated Content— source for Article 50 applicability from 2 August 2026, the Code's finalisation on 10 June 2026, the Commission's adequacy confirmation, the roughly 190 signatories by late July 2026, and the machine-readable marking requirement.
  4. EU Artificial Intelligence Act — Transparency rules, Article 50— the text of the transparency obligations for providers and deployers of generative AI systems.
  5. European Commission — Transparency obligations under Article 50 of the AI Act (FAQ)— source for the 2 December 2026 transition for systems placed on the market before 2 August 2026, the Article 50(2) exemption for assistive standard editing, the Article 50(4) deployer labelling duty and its human review and editorial responsibility exception, and the penalty range.
  6. Dathathri et al. — Scalable watermarking for identifying large language model outputs, Nature (2024)— the peer-reviewed paper describing SynthID-Text, which Anthropic names as the basis of its approach.
  7. Google DeepMind — SynthID— Google's description of SynthID across image, video, audio and text, including the statement that for text it adjusts probability scores to generate a watermark.
  8. Coalition for Content Provenance and Authenticity (C2PA)— the open standard Anthropic uses for signed provenance metadata on supported generated files.
Prepared by ADAM Pulse (USA Telecom Consulting LLC)

Managed network and communications services, SDVOSB. We are asked about AI governance mostly by clients who need a policy that survives an audit rather than a philosophy. If you need the four-way distinction written into something enforceable, we can help. Support: (888) 989-4872 · support@adampulse.us