Back to Blog
TIA™August 15, 20267 min read

The Watermark Is Not the Threat

On August 2nd, Anthropic started marking every word Claude writes. Not a tag on the file. Not hidden characters you could find and delete. The words themselves.

Within two weeks, the loudest question in my industry was how to get it out.

That question is the story. Not the watermark.

What it actually is

When a model writes, it is choosing each word from a set of candidates. Often several are equally good, and the choice between them carries no meaning. The watermark biases those coin flips against a secret key. Every individual word looks completely normal, because every individual word is completely normal. But across a few hundred words, the pattern of choices adds up to a signature.

This is not a theory someone floated in a blog post. It was published in Nature, and tested live across roughly twenty million Gemini responses to confirm it does not degrade the writing. Anthropic adopted the same approach and turned it on across everything: the API, the apps, the coding tools. Google has been doing it to images since 2023.

Two consequences fall directly out of the mechanism, and almost every take I read got both wrong.

It survives copy and paste, because you are copying the words and the words are the mark. And it dies under heavy rewriting, because you have replaced enough of the choices that the pattern falls apart. Light editing will not do it. Rewriting the piece will.

The second consequence matters more. A detected mark does not prove a machine wrote anything. This is Anthropic's own language: it cannot distinguish "Claude wrote this" from "Claude edited this." Write an essay by hand, paste it in, ask for a grammar pass, and the output carries the mark. Ask for a translation. Ask for a summary. Same result.

So the mark tells you one thing: this text passed through a model at some point. That is a much smaller claim than the one everyone is reacting to.

The chain everyone is running

The fear goes like this. Anthropic ships a detector. Google gets it. Google can now sort the entire index into human and machine. Google demotes the machine half. Every page you published becomes a liability overnight.

Every link sounds reasonable. The second one has never survived contact with evidence.

Google has been able to detect its own AI-generated images since 2023. Its own models made them. Its own marks are in them. Free detectors flag them in two seconds with no login and no key. Three years, complete capability, zero ambiguity.

It has never once been what decides whether a page ranks.

Capability and use are not the same thing, and the gap between them has been sitting in public for three years while the industry assumed it was closed. Google's published position is that appropriate use of AI is not against its guidelines, and that what matters is whether the work is original and useful. You should not take that on faith. Google has said things about ranking that turned out to be false under oath. But you do not need faith here, because you have three years of behavior, and the behavior says the same thing the blog post does.

The advice that will actually cost you

Someone is going to tell you to strip the mark anyway, just to be safe. The methods are real and published. Run it through a paraphraser. Round-trip it through another language. Swap characters for identical-looking ones from other alphabets.

The first two work, and they cost you the writing. You cannot use another major model to do the paraphrasing, because they are signing on to the same requirement under the same European law. So it is by hand, on a page that was already working, and the prose flattens and errors creep in that nobody proofreads for.

The character swaps are worse, and this is the part people do not see coming.

Google does not match keywords. It resolves entities. It reads a city name and connects it to everything it knows about that city. It reads your business name and connects it to your listing. None of that happens visually, because the algorithm does not have eyes. Text becomes tokens before anything looks at it. Replace a Latin a with a Cyrillic one that renders identically and you have not created a typo it can work out from context. You have created a token that resolves to nothing.

Do that to your service name, your city, your company, and the page stops making the connection it was built to make. Mixed-script text is also trivially detectable, and has been a spam signal for years.

You would be breaking the meaning of your own page, handing over something that actually is flagged, to hide from a consequence that has not happened.

The part nobody is saying

Here is what convinced me the panic is not about search rankings.

The mark cannot tell writing from editing. In December, the transition period ends for models that were already on the market, and the rest of the industry has to meet the same requirement. Follow that out. Within a year the mark is on the essay written entirely by a machine, on the essay a person wrote and ran a grammar check over, and on the memo someone dictated and had cleaned up.

A signal that is present on nearly everything carries almost no information.

So the thing people are scrambling to remove is, on its own, close to meaningless. Which means the scramble is not about the signal. It is about what the signal being available implies: that the question could now be asked at all.

That is a confession. It says the plan was to be indistinguishable, and being indistinguishable was load-bearing.

I have never hidden that I use AI. I do not shout it from the rooftops either. It is simply how the work gets made, the way a spreadsheet is how the math gets done, and I have never once needed a reader not to know. So when the mark arrived, nothing about my position changed, because my position never rested on the thing the mark reveals.

That is not a moral claim. I am not interested in scolding anyone about disclosure. It is a structural one, and it is the actual lesson.

What this is really about

Any strategy that depends on a capability you do not control is not a strategy. It is a bet on the world staying still.

Undetectability was never a capability you owned. It was a temporary property of somebody else's product decisions, and it held right up until two engineers changed a sampling procedure. Everyone whose position rested on it lost that position in a single announcement, with no warning and no recourse, and the only move left was to degrade their own pages trying to claw it back.

The people who did not lose anything are the ones whose work was defensible when it was known. Not because they are more honest. Because they built on something that could not be taken away by a release note.

So the question worth asking is not how to remove the mark. It is this: if every reader knew exactly how your work was made, what would change?

If the answer is nothing, you can stop reading about watermarks.

If the answer is everything, the watermark was never your problem.


Sources: How Claude's text watermark works — Anthropic · How Claude marks AI-generated content — Anthropic · Scalable watermarking for identifying large language model outputs — Nature

Jon Mayo

Written by

Jon Mayo

Liked “The Watermark Is Not the Threat”?

Get notified when new TIA™ articles are ready.

Subscribed to: