On October 5, 2026, OpenAI said it would start injecting invisible watermarks into text produced by ChatGPT and Codex for users in the European Union. The rollout, expected "over the coming weeks," is the company's answer to the EU AI Act's requirement that providers label machine-generated text in a machine-readable way. It is a genuine compliance milestone. It is also a revealing test of how far text watermarking has actually come.
The technology carries a name with weight behind it: textGrain. Instead of stamping text with hidden characters, it embeds an invisible statistical signal in the model's choice of words. OpenAI says textGrain matched or exceeded other approaches, including Google DeepMind's SynthID for text, which itself is the basis of the watermarking Anthropic announced in August. In the company's own tests, with the detector set to a false-positive rate of one percent, detection of a 400-token passage about psychology landed around 95 percent; at 200 tokens that fell to roughly 80 percent. OpenAI also reported no significant quality differences across eight benchmarks when the watermark was switched on or off, using its frontier model Astra.
The Claim
On paper, this reads like the transparency breakthrough regulators hoped for. Watermarking is on by default. It does not change the words a reader sees. Copy and paste preserve the pattern. OpenAI says it plans to release the technology as open source so others can build on it. Every one of those points is true.
But the most important number in the announcement is the one the marketing does not put first. Editing defeats the watermark quickly. Replace just ten percent of the words with synonyms, and detection for 400-token passages drops from about 92 percent to 66 percent. Replace a quarter of the words and detection collapses to 17 percent, which is effectively undetectable. For mathematics, where the model has less freedom over word choice, the rates were substantially lower. Short text suffers the same way.
In other words, a person who simply rewrites one out of every four words walks away clean. That is not a corner case. It is how anyone edits an email, a report, an essay, or marketing copy. The watermark survives copying but not thinking, which is precisely the moment it is supposed to reveal itself.
A Compliance Checkbox More Than a Guarantee
The second thing worth noticing is where this applies. Watermarking turns on by default only in the EU. API customers everywhere else may opt in, and OpenAI made that opt-in available worldwide, in contrast with Anthropic, which makes watermarking mandatory for Claude no matter where or how it is used.
OpenAI frames the regional gate as deliberate. A regional approach, the company says, gives room to learn from real-world use and feedback. Read skeptically, that is an honest admission that the system is early and unproven, which is exactly why it is being rolled out as an experiment in the one jurisdiction that demands it rather than everywhere at once. Transparency here is a tax paid where it is required, not a feature shipped because it is good.
What the Watermark Cannot Tell You
OpenAI is candid that a detected watermark reveals nothing about how much human creativity or editing shaped the text. It does not establish ownership, assign responsibility, or separate a correct answer from a wrong one. And failing to detect a watermark does not prove a human wrote the passage; the text may simply have been short, edited, translated, or produced by a model the detector does not cover.
What the watermark genuinely cannot tell you:
- that a human or an AI wrote the text, since a missed watermark proves nothing about the author
- who is responsible for the content, because the signal carries no identity or ownership claim
- whether the information is accurate, because watermarking says nothing about truth
- that the output is unedited, since editing and translation are exactly the moves that break it
The detector itself is not available to the public. OpenAI is granting access case by case to selected researchers and specialist organizations under the EU Code of Practice, and the tool reports only whether an OpenAI watermark was found, without revealing users or their prompts. The company says it will widen access when results can be interpreted responsibly. Until that happens, the people being asked to trust the watermark cannot verify the claims that justify it.
None of this means textGrain is worthless. Invisible statistical watermarking is genuinely hard, and OpenAI's honesty about its failure modes is refreshing in an industry that usually ships a slider and calls it safety. The problem is the framing around it. If the EU AI Act is read as machine-readable labeling, the honest engineering answer is that the label can be wiped out by a modest edit, and the number of parties that can actually read it today is a small, hand-picked group.
The practical takeaway for anyone trying to spot AI text is simple. Treat the watermark as a weak prior, not a verdict. A detected watermark is meaningful evidence; the absence of one is nearly meaningless, because the text may have been rewritten, translated, or generated by a model that does not participate. And the moment a quarter of the words change under a human hand, the signal is gone.
So welcome the compliance win, but do not overstate the product. OpenAI is doing something real, in a narrow place, with an early technology that a 25 percent rewrite can defeat. That is worth remembering before anyone tells you AI text is now reliably identifiable.
Comments