# Why I stand against AI watermarking

On August 11, Anthropic published [How Claude marks AI-generated content](https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content): the official account of the invisible ink it is putting in Claude's sentences. A statistical fingerprint woven into the words.
Invisible to readers. Applied worldwide, not just in the EU. Built to satisfy the [EU AI Act](https://eur-lex.europa.eu/eli/reg/2024/1689/oj)'s transparency obligations (Article 50). And, they promised, detectable not only by Anthropic but by _"users and other third parties."_
I read the article and felt something tighten in my chest. Not because I hate regulation. Because I knew, in that moment, that the mechanism was wrong, technically, legally, ethically, and that somebody should say something about it. The solution, it turns out, was in their own help center.
Four hours later, I shipped a tool that attempts to remove those watermarks. The internet did the rest: 2M impressions, 40k+ engagements, +9k stars on GitHub, dozens of contributions, in less than three days. Commenters called me everything from a hero to a regulatory saboteur, some even generating AI images of me... [[1](https://x.com/PinkSilkPham/status/2088238182072746125)], [[2](https://x.com/krishdotdev/status/2087983118544474356)], [[3](https://x.com/guillaumemeyer/status/2087974300246819080)].
I'm neither. I built [watermarks-remover](https://github.com/guillaumemeyer/watermarks-remover) because:
1. I am a nerd, and I wanted to see how it was done.
2. I'd rather argue about this with my hands dirty than from the sidelines.
Let's make it clear up front:
> I'm not against content attribution. I'm against the watermarking technique.
Everything that follows argues against the technique, not against attribution itself.
### The watermark answers a question nobody asked
> A text watermark tells you exactly one thing: _a machine was involved._
That's it.
Anthropic says so themselves: A detected mark _"indicates that the content may have been processed by Claude."_ It _"is not fully conclusive."_ It _"does not, on its own, confirm the full provenance of the content."_
They even list the cases that break their own binary. _"Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source."_ That is the entire authorship problem, in their words, on their support page.
> Authorship is a spectrum, not a binary.
The watermark treats it as a binary (not necessarily technically, but practically speaking). It stamps the person who did 95% of the work the same way it stamps the person who did none.
And here's the cruelest part, also theirs: a watermark dies the more you edit. They warn that a mark may vanish if _"the text has been heavily edited, paraphrased, translated, or mixed into other writing,"_ or if _"the passage is very short, leaving too little text for a reliable signal."_ The more you make a text your own, the more you rewrite, restructure, translate, argue with the model and win, the faster the fingerprint degrades into noise.
Think about what that means. The watermark is most reliable precisely when the human contributed the least. It is least reliable when the human contributed the most.
> The system punishes exactly the people doing the most authorship work, and rewards, with invisibility, the people who did none.
That's not a flaw in the implementation. That's the design.
### What the mark actually claims
Let me be precise, because precision matters here, and because Anthropic asked for it. A watermark is not a copyright claim.
> A watermark doesn't say "Anthropic owns this." It says "a machine was involved."
Those are very different statements. The danger is the slide between them. Once you accept that "a machine was involved" means "a machine wrote this," it's a short walk to "the vendor's machine wrote this, so the vendor has a claim on it."
Conflating involvement with authorship erases the human's contribution and hands the moral weight of our words to model vendors.
I didn't sign up for that. Nobody did.
### It's not reliable enough to be evidence, and it's not harmless
Here's what the help center says: detection is statistical. A positive result is a verdict with a confidence score: a probability, not a fact: ["A detected mark provides a signal that content was processed by Claude, but is not fully conclusive."](https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content)
The inverse is worse. ["Lack of a detected mark doesn't mean the content wasn't AI-generated or processed."](https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content) So a hit isn't proof, and a miss isn't clearance. Every detector trades misses against false positives. Tune it to catch more watermarked text, and you accuse more innocent people.
And they are not keeping that engine in-house. The same article promises to ["support users and other third parties to detect Claude's marks."](https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content) That means the probabilistic accusation engine gets handed to schools, employers, platforms, and anyone else who wants a dashboard.
> Every false positive will have a human face attached.
A student accused of cheating on a paper she wrote. A writer accused of faking an article she agonized over. An employee accused of fabricating a report she drafted, edited, and owned.
### It does not match Anthropic's own values
Anthropic built its reputation on AI safety. Its entire brand is _"we are the ones who think about harm."_
> Shipping a probabilistic accusation engine, and documenting in public that it is not conclusive, is not safety.
It is a new way to harm innocent people, deployed by the one company that promised, louder than anyone, that it would never do that. Judged by their own stated values, this fails. The tool meant to protect people from AI harm may become a source of it.
### The stakes aren't hypothetical. Ask your doctor.
This is not a hypothetical future. At Kaiser Permanente alone, [7,260 physicians used generative AI scribes across 2.5 million encounters in a little over a year](https://catalyst.nejm.org/doi/full/10.1056/CAT.25.0040). The doctor reviews, corrects, and signs the note.
> The doctor is the author. The doctor bears the responsibility: legal, ethical, professional.
Now stamp that note "machine-generated." Patients start doubting the note, the diagnosis, the doctor. Trust in the medical record (the thing patients rely on when they're at their most vulnerable) erodes over a label that, in Anthropic's own phrasing, says only that a machine "may have been processed" into the text.
I can already hear the counterargument: _shouldn't patients know a tool was involved?_ Yes: they are, as they have to accept the recording of their visit. But patients should know the doctor is responsible. That's the label that matters. A watermark doesn't say who is accountable; it says a machine touched the text, and it lets everyone draw the wrong conclusion.
> Mislabeling an accountable clinician's record isn't transparency. It's damage.
### Who stays marked?
Here's the asymmetry nobody mentioned, and that the support article quietly concedes. A watermark only survives untouched text. Anthropic lists the escape hatches themselves: heavy editing, paraphrasing, translation, mixing the output into other writing. Anyone with basic knowledge, including a tool like [watermarks-remover](https://github.com/guillaumemeyer/watermarks-remover), can strip one in seconds.
So who stays marked? The honest user. The student who doesn't know the fingerprint is there. The blogger who never heard of a detector. The disinformation operator, the scammer, the fraudster? Unaffected. They'll strip the mark before they hit publish.
> Watermarking text, as deployed, is a tax on the naive.
It catches the people who never needed catching and lets the people it was supposedly designed for walk right through. The EU wanted transparency against deception. It got a label that only slows down the people who were never the problem.
### What I'd do instead
Let me be unambiguous, because I keep being cast as the villain of a movie I'm not in: I am not against transparency. I'm against the one method that's guaranteed to fail.
The irony is that Anthropic already knows the better method. The same article describes a second marking system for files: signed provenance metadata, following the C2PA open standard, attached to generated files. That is the right idea: a label at the container, visible to any C2PA-aware tool, honest about what it is.
Transparency should be built where it survives:
- **Provenance at the platform level.** Metadata, content credentials (the C2PA-style standard they already use for files) attached at the moment of generation, where they can't be silently confused with authorship.
- **Watermarks as at most a weak signal.** A hint, never a verdict. Never evidence. Never the basis for accusing a student, a writer, a doctor. Their own article already says a detected mark is "not fully conclusive." Treat it that way.
- **Honest labels.** "This was AI-assisted," said plainly, at the point of sharing, by the person doing the sharing. A label, not a hidden fingerprint.
> That's the standard I'd hold any lab to: transparency that survives editing, that respects the author, and that never mistakes a probability for a fact.
### The signal
People keep asking me why the tool went viral. I think they're asking the wrong question. The virality wasn't about me. I was just the person who happened to ship first.
It was a vote. Dozens of thousands of people, in English, French, Spanish, Italian, German, Portuguese, Russian, Chinese, Japanese, Korean, in less than 48 hours, independently deciding they don't want their own drafts marked.
That's the signal regulators and labs should be reading. When you stamp invisible ink on people's words, you find out how many people consider their words their own. The answer, it turns out, is almost all of them.
> Machine-involved is not machine-authored, our words are ours, and we deserve better than a fingerprint we never agreed to carry.