Anthropic turns on default AI watermarks after EU law, but buyers can’t verify
An arXiv preprint says Article 50 of the EU AI Act took effect on August 2, 2026, and days later Anthropic said Claude models would watermark text by default using SynthID-Text; Google’s Gemini has used SynthID-Text since 2024. The paper argues the real problem for enterprises is unverifiability: ne
Hannah Vogel ·

In an arXiv preprint dated September 2026, the authors say Article 50 of the EU AI Act took effect on August 2, 2026, obliging generative AI providers to “mark the content their systems produce and ensure it can be detected as AI-generated.” Days later, they write, Anthropic disclosed that “every Claude model released after that date embeds a watermark based on SynthID-Text in all generated text, enabled by default with no user opt-out,” and that Google has deployed SynthID-Text in Gemini since 2024. This is, so far, single-source — an arXiv preprint — with no independent confirmation or on-the-record interviews in the packet. The paper’s core claim is not about the watermarking technology itself, but about governance: neither the vendors’ assurances nor users’ objections can currently be verified, and that unverifiability is the actual compliance risk for enterprises operating under EU rules. See: “Watermarks Without Verification: AI Text Watermarking After the EU AI Act,” arXiv (http://arxiv.org/abs/2609.09604v1).
The law moves the detection burden to providers; the proof problem lands on buyers
According to the preprint, Article 50 places two linked obligations on model providers: mark outputs and ensure they are detectable as AI-generated. That framing sounds like providers carry the load. In practice, enterprises that publish AI-assisted content and code will be asked to demonstrate that their vendors meet the detection standard, or at minimum to show they can distinguish AI-origin content when regulators, platforms, or counterparties ask. The authors report that “no public tool can test the deployed systems.” They therefore evaluate the open-source SynthID-Text implementation on two open-weight models, and find that on prose, the measured quality effect of watermarking “does not exceed that of changing the sampling seed.” On code, they measure a three-point correctness hit on one model and “below measurement” impact on the other, while detection “remains near chance,” which they attribute to detectability limits rather than a quality problem. The preprint cautions that these findings are not tests of Anthropic’s or Google’s proprietary deployments and cannot verify their claims — and that is the point. Without vendor-disclosed evidence, buyers cannot close the proof gap.
Default watermarking with no opt-out is a contract and SLA issue, not a feature toggle
If Anthropic has indeed enabled watermarking by default with no user opt-out, as the preprint asserts, the control point moves out of the prompt and into the vendor’s configuration and policies. That makes it less a product setting and more a service-level and indemnity matter. For marketing and communications teams, the change could be benign — the paper’s prose result implies little to no quality degradation in copy. For engineering teams using code generation, the measured three-point dip on one open-weight model raises a different question: not whether a specific vendor incurs the same cost, but whether a team can even know. The preprint maps what would be needed to settle such questions: release of matched outputs, configuration disclosure, accredited audits, a shared evaluation protocol, and interoperable detection. Until those exist in contracts and service documentation, procurement cannot rely on slideware assurances, and legal cannot confidently represent compliance.
Vendors say quality is unchanged and no identifiers are embedded; critics say the opposite — neither side can be tested today
The preprint catalogs objections from users — that watermarking degrades quality (especially for code), “secretly encodes identifying information,” and is both “easily removable and inescapable” — alongside vendor assurances of “unchanged quality,” “no identifying information,” and “robustness to light editing.” It argues these claims are mutually contradictory and, more importantly, unverifiable at present. Because there is “no public tool” for the deployed systems and model providers have not released matched outputs or configurations, independent labs cannot replicate vendor results or falsify user complaints. For enterprise buyers this is not an academic quibble: a dispute about detectability or hidden identifiers becomes a privacy, IP, or consumer-disclosure risk if contested by a regulator or platform. The preprint’s evaluation on open-weight models suggests that, in code, detectability can be near chance even when a watermark is present, which is a detectability failure that no amount of vendor marketing can paper over without evidence.
Marketing and code teams face different costs — and different proof requirements
If the preprint’s findings generalize, brand and PR teams should expect negligible copyquality impact from watermarking relative to normal sampling variance. Their problem is proof: will their CMS or distribution platforms accept a vendor’s claim that text is watermarked, or will they require demonstrable detection under a shared protocol? Agencies that ghostwrite for multiple clients will want interoperable detectors so one tool can test content across vendors. Engineering leaders face a more operational trade-off. A small-but-real hit to code correctness on some models is manageable if detected early and if teams can turn it off where unacceptable. But in a “default on, no opt-out” world, mitigation shifts from a project setting to vendor selection and contract carveouts. If the watermark is robust to light editing, as vendors claim, post-generation refactoring may not clear a detector — which means CI/CD or code-review processes might need to accommodate detection noise until tools improve, or teams shift to providers with audited performance on code.
Interoperability and accredited audits are the missing institutions; without them, compliance stays a spreadsheet
The preprint is most useful where it stops pretending to adjudicate technology claims and instead enumerates institutional gaps. It calls for: release of matched outputs so that independent testers can measure detection accuracy without reverse engineering; disclosure of watermark configuration so that robustness claims are falsifiable; accredited audits of watermarking and detection so enterprise buyers can rely on third-party reports rather than vendor assurances; a shared evaluation protocol so regulators, platforms, and enterprises test against the same yardstick; and interoperable detection across vendors so downstream distributors can accept a single attestation. These are procurement artifacts as much as research desiderata. They can be written into RFPs, DPAs, and SLAs. Providers that meet them first gain an immediate commercial advantage because they reduce the buyer’s compliance workload.
The skeptic’s read: proprietary deployments may be better — let them prove it
A fair counter is that open-source implementations on open-weight models are not representative; proprietary systems may incorporate watermarking that is both higher-quality and more detectable. The preprint effectively invites that rebuttal and lays out what evidence would settle it. If a major provider publishes an accredited audit complete with matched outputs, configuration ranges, and cross-vendor detection results, the claim of superior performance becomes verifiable and actionable in procurement. If, instead, providers continue to assert “unchanged quality” and “robust to light editing” without artifacts, the compliance burden stays on the buyer to reconcile contradictions under deadline. Under EU rules as described by the preprint, “ensure detectability” will need teeth — either in regulatory guidance with measurable thresholds or in market practice via shared protocols. Until then, expect longer renewals, heavier legal review, and a premium for vendors willing to be tested.
What changes this quarter for buyers and sellers
For buyers: add watermarking evidence to security and privacy questionnaires; ask for matched outputs for specified prompts, configuration disclosures, and an audit timeline; and price the risk that detection remains near chance for code or introduces false positives in copy workflows. For sellers: decide whether to offer detection APIs and audit-ready artifacts as paid support or as table stakes; clarify whether watermarking is configurable per use case; and prepare for agencies and systems integrators to ask for cross-vendor interoperability commitments. The preprint’s central warning is that unverifiability is the governance failure. In enterprise terms, that means unverifiable claims do not belong in compliance checklists — they belong in the contract, with tests you can run before you sign.