James Padolsey's Blog

2026-08-12

Anthropic’s weak watermarks appease a weak law

Anthropic announced in August 2026 that text generated by supported Claude models will carry invisible, machine-readable watermarks, applied at model level across supported Claude products and API surfaces worldwide. The move is Anthropic’s response to Article 50(2) of the EU AI Act, which requires providers to ensure that outputs are “marked in a machine-readable format and detectable as artificially generated or manipulated”. Yet the same provision exempts systems performing an “assistive function for standard editing”, or which do not substantially alter the user’s input or its meaning. This is a provider-side technical obligation, not a general requirement that every person using AI must visibly disclose it.

Anthropic concedes that the mark is not proof of authorship: Claude may merely have proofread, translated, summarised or otherwise processed human work, while substantial rewriting, paraphrasing, translation or mixing may make the mark disappear. It may survive copy-pasting and light editing. It is a hint of provenance, not a definitive detector.

I’m not against such provenance measures, nor against ensuring humans remain accountable to each other. And I don’t want cognitive decline due to AI replacing our prose-writing. However, I am quite against appeasing laws that are not truly efficacious, or rather, are efficacious only against people who are not technically sophisticated enough to circumvent them. Not only that, but in this case that is likely to correlate with those who have a higher need for AI-aided text. I wager this disproportionately includes disabled and neurodivergent individuals, or those in circumstances that afford them less time and capacity for the ordinary expenditure of cognitive effort involved in prose-writing.

That is to say: this rule risks penalising people who are already less able to produce conventional prose unaided for using technological means to make their lives easier. The Act exempts standard editing, but many legitimate assistive uses require more substantial rewriting while leaving the ideas, judgment and responsibility with the human. The same thought that led to this law could have applied to calculators at the time of their inception, had their outputs revealed themselves through artefacts. Thankfully, a sum borne of the brain is treated no differently from one produced by a calculator. Likewise with spellcheckers. To make assistance suspect only once the tool becomes capable enough to compose a whole sentence is not a principled boundary. It is a moral premium placed on difficulty itself.

Anthropic has nevertheless chosen a blanket, model-level implementation that appears broader than the law’s minimum requirement. That may be convenient compliance engineering, but it discards distinctions the law expressly attempted to preserve. The result is a signal broad enough to implicate harmless and assistive use, yet fragile enough to be removed by a motivated person through substantial recomposition. It risks concentrating suspicion on ordinary and assistive users while remaining weakest against deliberate deception.

As an engineer I feel beholden to enable people to use technology in ways that serve harmless purposes. A weak watermark may still be useful against verbatim or lightly edited model output. But it is weakest precisely against motivated users prepared to remove it. Naturally this isn’t about Claude, is it? It is a reaction to a technological scare, a reactionary provision that, unlike much of the mostly reasonable EU AI Act, lacks technical prowess and fair purpose.

Anthropic is, primarily, in the business of protecting Anthropic. In my own tests, I noticed this directly when I pressed Claude on how text watermarks work and how recomposition can disrupt them. Asked in general terms, or about a competitor’s watermark such as Moonshot/Kimi, it reasoned freely. But when the subject turned to Anthropic’s own mark, the same general questions drew visibly more caution and reluctance to help. This does not establish why the model behaved differently. But the asymmetry is at least consistent with self-protection rather than a stable, provider-neutral principle.

Screenshot of declaude.org: sycophantic Claude-flavoured text pasted on the left, the plain-prose rewrite on the right, with the app reporting 3 em-dashes removed and 7.7% of the original phrasing surviving

I hope, in remedy to this, people will make use of declaude.org, a little web app that obscures and remixes AI-generated text in ways designed to disrupt various types of fingerprinting, including subtle context-based token-sampling biases. Anthropic has not yet released its detector or full technical documentation, so I cannot yet claim that declaude defeats Claude’s particular watermark. Nor does it remove any genuine legal, academic or professional duty to disclose AI assistance. But a watermark that burdens the candid and yields to the deceptive is not meaningful transparency. It is compliance theatre.


By James.


Thanks for reading! :]