Key takeaways
- OpenAI’s textGrain embeds an invisible statistical watermark in ChatGPT and Codex text for EU users, rolling out over the coming weeks, while API customers worldwide can opt in for select models. It stays off by default outside the EU. The detector goes only to approved researchers and expert organizations, case by case.
- The mechanism is entropy-calibrated sampling: a secret key and the preceding words steer token choice through an optimal-transport coupling, spending a capped “entropy budget” so the text reads the same and benchmark scores barely move. The detector needs only the text and the key.
- OpenAI published its own failure modes next to the launch. At a 1% false-positive target: about 80% detection on 200-token passages, 95% on 400-token ones, substantially lower on mathematics. Swapping 10% of words for synonyms took 400-token detection from about 92% to 66%; 25% took it to 17%.
- A watermark is evidence when you find it and nothing when you don’t. OpenAI’s own line: the absence of a detected watermark does not prove human authorship. If your compliance plan treats the mark as a provenance log, an authorship label, or a truth check, rewrite the plan.
On October 5, OpenAI published a post with an unusually honest shape for a product announcement: a new capability, followed immediately by the company’s own measurements of where it breaks. The capability is textGrain, a text watermarking system for ChatGPT and Codex. The reason is the EU AI Act. The breakage is the interesting part, and OpenAI printed it in the same post.
Article 50(2) of the AI Act requires providers of generative AI systems to mark synthetic text, audio, image, and video outputs “in a machine-readable format,” detectable as artificially generated. The transparency duties went live on August 2, 2026. Systems already on the EU market before that date have until December 2, 2026 to add the marking. ChatGPT and Codex were on the market long before August, which is why OpenAI says the watermark arrives “over the coming weeks” rather than on a date.
This is the first time a European AI rule has produced a different version of a mainstream AI product for EU users. That alone makes it worth understanding, even if you never serve a single EU user. The mechanism is genuinely clever, the limits are genuinely limiting, and the gap between the two is where the compliance homework lives.
What OpenAI actually shipped
Three tracks, announced together.
First, the EU default. Over the coming weeks, eligible text output from ChatGPT and Codex gets an invisible watermark for users in the European Union, across all plans. OpenAI is explicit that it is not making text watermarking a global default at launch: “This regional approach gives us room to learn from real-world use and feedback.”
Second, the API opt-in. Starting October 5, API customers anywhere can switch on watermarked text output for select models. It stays off by default, it is a setting rather than a per-request parameter (organization or project settings), and OpenAI says it is working with cloud partners so models served through them can be watermarked too.
Third, the restricted detector. Approved researchers and expert organizations can apply for access, granted case by case. The tool reports only whether it detects an OpenAI watermark. It does not identify the user and does not reveal prompts or conversations. OpenAI also says it plans to release the technology as open source so other labs can build on it, without committing to a date.
Note what is not in the announcement: a public checker. You cannot paste a paragraph anywhere and ask. Unless your team is an approved research or expert organization, nobody on it can verify a watermark either way. Keep that in mind for everything that follows.
How textGrain works, in plain terms
A language model writes one token at a time, a word or part of a word. At each step it holds a probability distribution over what could come next, then rolls the dice. A text watermark replaces part of that dice roll with pseudorandom numbers derived from a secret key and the words that came before.
textGrain’s version, detailed in a technical report that lays out the coupling math, co-written with researchers from the University of Pennsylvania and Yale, works like this. The secret key and a window of preceding tokens divide the vocabulary into blocks. An optimal-transport coupling then nudges the sampler toward whichever block the key favors at that position. It spends from an “entropy budget”: a cap on how much of the model’s natural randomness may be traded for watermark signal. Inside the chosen block, tokens keep their original relative probabilities, so the model still picks a sensible word.
The report’s own example makes it concrete. After “The morning was,” the model might give “warm” 30%, “cold” 25%, “mild” 15%, and “calm,” “sunny,” and “bright” 10% each. Any of those continues the sentence fine. textGrain uses the key to favor one block of candidates over another. Across hundreds of such nudges, a statistical pattern accumulates. The detector, holding the same key, tests for it. The detector needs only the text and the key. No model, no prompt, no generation log.
Two design choices matter. First, the watermark is unbiased: averaged over all possible keys, the model’s word choices are unchanged. Only the source of randomness changed, not the distribution. That is why OpenAI reports no meaningful benchmark movement with watermarking on. Second, the entropy budget fixes a real flaw in older schemes. The classic Gumbel-max watermark always picked the same token for the same context and key, so asking the same question twice could produce byte-identical answers. Spending only part of the randomness keeps repeat answers different, which matters when an application needs several distinct responses to one prompt.
The idea is not new, and the report does not pretend otherwise. It cites the 2023 line of work that started the field, including Kirchenbauer and colleagues’ green-list watermark, which biases sampling toward a pseudorandom subset of tokens and detects the surplus with a statistical test. textGrain is a more careful instrument in the same family: calibrated, budgeted, and honest about its limits.
Strip it yourself
Reading about a statistical watermark is one thing. Watching one die under your own edits is another. The lab below runs a teaching model of this watermark family on a fixed passage. Swap synonyms, run a translation round-trip, or switch to constrained math-like text, and watch the detector’s p-value move. Then read what the verdict can and cannot prove.
Interactive: the attack lab
This is a teaching model of the watermark family textGrain belongs to, not textGrain itself. A fixed passage below carries a keyed statistical watermark, generated the same way the demo detector expects. Swap synonyms, run a translation round-trip, or switch to constrained math-like text, and watch the detector's p-value collapse. The decay matches the shape of OpenAI's published figures.
Calibration: OpenAI's own tests, at a 1% false-positive target, found the watermark in about 80% of 200-token and 95% of 400-token psychology passages. Replacing 10% of words with synonyms took 400-token detection from about 92% to 66%; 25% took it to 17%. This toy reproduces that collapse on a short passage.
word swapped word
Detector p-value: -
Detection threshold 0.01: below it, the detector calls the watermark present.
- words scored
Three things to notice if you ran it. First, the collapse is nonlinear. A few swaps weaken the signal; at around a tenth of the words rewritten, the detector starts to fail; past a quarter, it is gone. That mirrors OpenAI’s published curve, and it is why “lightly edited” is the most dangerous phrase in any watermarking policy. Second, the math-like passage fails before you touch it. Constrained wording leaves the sampler almost no freedom, so there is almost no signal to find. Code, formulas, names, and dates are the same story. Third, shortening the passage weakens detection even with zero edits. Signal accumulates per token; short text simply has less of it.
The numbers, honestly
All figures below are OpenAI’s own, at a target false-positive rate of 1%, which means about one unwatermarked passage in a hundred can come back positive by design.
| Condition | Detection rate |
|---|---|
| 200-token passage, flexible prose (psychology) | About 80% |
| 400-token passage, flexible prose | About 95% |
| 400-token passage, mathematics | Substantially lower |
| 400-token passage, 10% of words swapped for synonyms | About 92% to 66% |
| 400-token passage, 25% of words swapped | About 17% |
On output quality, OpenAI’s benchmark table for Astra, its latest frontier model, moves in both directions with watermarking on:
| Benchmark | Unwatermarked | Watermarked |
|---|---|---|
| Artificial Analysis Intelligence Index | 49.57 | 49.76 |
| AutomationBench | 34.09% | 34.86% |
| GPQA Diamond | 94.44% | 93.94% |
| Terminal-Bench 4.0 | 53.90% | 56.06% |
| Terminal-Bench Science 0.1 | 56.90% | 60.00% |
| BrowseComp | 87.92% | 87.35% |
| DeepSWE v1.1 | 72.80% | 71.68% |
| HealthBench Professional | 64.27% | 64.60% |
No consistent decline, which is what the unbiasedness property predicts. Practitioners on Reddit remain skeptical that watermarking leaves quality untouched, with some fearing a “constrained version” of output; the published numbers are the counterweight, and they are OpenAI’s own tests, not independent ones.
Two more honest details. The watermark adds no hidden characters, no invisible spaces, no unusual punctuation: the signal lives entirely in which words were chosen, so a character cleaner has nothing to clean. And the mark travels with copy and paste, because it is the wording, not metadata. That is the good news for durability. The bad news is everything in the table above.
Five things the watermark cannot tell you
OpenAI devotes a full section of the post to this, and it reads like it was written by someone who has watched detection results get misused. Paraphrasing their five points:
- It does not measure human contribution. It can indicate that an OpenAI system generated or processed part of a passage, but not how much human judgment, editing, or creativity went into it.
- It does not establish ownership or responsibility. It says nothing about who owns the text, whether its use was lawful, or who answers for it.
- It does not identify the user. No person, organization, account, prompt, or conversation is associated with the mark.
- It does not verify accuracy. It says nothing about whether the passage is true, misleading, or harmful.
- Its absence proves nothing. Text may be too short, edited, translated, from an unsupported model, predating watermarking, or from another company’s tools.
Point 5 is the one to print out. Practitioner discussion since the announcement keeps circling the same objections: the mark is fragile against paraphrasing and translation, the EU-only rollout looks like minimal-effort compliance, only the provider holds the detection key, and a missing mark will inevitably be read as proof of human authorship by people who should know better. The first three are engineering and policy facts. The fourth is a communication failure waiting to happen, and it is the one your internal guidance can actually prevent. Write the sentence “a negative result does not prove a human wrote this” into every policy that mentions watermarking, because someone will need to read it back to an executive.
There is also a sharper privacy angle worth naming. Detection is asymmetric by design: only the key holder can check. If a third party ever offers “was this written by AI” screening built on a provider’s detector, every check means sending the text to someone’s servers. Centralized detection of decentralized text is a data-protection problem wearing a transparency costume. Treat any vendor selling that accordingly.
The compliance homework
Article 50(2) binds the provider of the generative AI system, not the person using it. If you build a product that writes text and put it on the market under your own name, that is very likely you: under the Act’s definitions, the company placing the system on the market is its provider, even when the model underneath comes from someone else. A switch nobody flipped is not a marking solution.
The mechanics that matter for the homework:
- The clock. Transparency duties have applied since August 2, 2026. Systems on the EU market before that date have until December 2, 2026 to add 50(2) marking. If your product serves EU users, that date is your deadline, not OpenAI’s.
- The floor. The Code of Practice sets no watermarking requirement under 200 tokens, roughly 150 English words. Short outputs are out of scope by design.
- The exemptions. Source code is out: the guidelines exclude code written to be interpreted, compiled, or executed. So are short outputs like single words and UI labels, agent-to-agent traffic no human sees, and standard editing like grammar fixes. AI-generated summaries, on the other hand, count as content needing marking.
- The penalty band. Breaches of the transparency duties can draw fines up to 15 million euros or 3% of global annual turnover, whichever is higher.
What to do on Monday, if any of your product’s text reaches EU users:
- List every place your product returns generated text to a user. For each, write down the provider, the model, the typical length, and the kind of output. Long, freely generated prose is what the rule targets.
- Make an explicit decision on OpenAI’s Text provenance switch per project, with a date. Off by default is not a decision; it is the absence of one. If you route between providers or models, remember that marking is a property of the route: the same prompt can produce marked text on one model and unmarked text on another or on an older version.
- Log which model produced each output. Your own log is no substitute for marking, but it is the only provenance record you can actually query yourself.
- Do not build anything that depends on detection. Not a plagiarism check, not an authorship filter. You cannot get the detector, and a negative result proves nothing.
Two vendors, two theories
The contrast with Anthropic is instructive. Anthropic signed the same Code of Practice and ships watermarking as a global default: every supported Claude model, every surface, everywhere Claude is offered, with C2PA signed metadata on supported generated image files and a detection API planned. OpenAI ships EU-only by default and opt-in elsewhere. One vendor treats the mark as a product property; the other treats it as a regional compliance control.
Which theory wins will show how far the Brussels effect reaches in AI. If other labs follow Anthropic, machine-readable marking becomes the default texture of AI text worldwide, and the EU rule will have set a global standard the way its data-protection law did. If they follow OpenAI, marking stays a regional overlay, and the interesting question becomes what happens at the boundary: how “EU user” is determined, what a VPN does to it, and whether API customers outside the EU ever flip a switch that stays off by default.
For Canadian teams, the jurisdictional note is short. The AI Act reaches providers placing systems on the EU market regardless of where they are established, and outputs used in the EU. A Toronto company serving EU users is in scope as a downstream provider. Canada has no equivalent rule: the federal AI transparency consultation closed on September 23, 2026, and whatever follows will take its own path. The homework above applies the moment your text crosses the Atlantic, whether or not Ottawa ever writes its own version.
Actionable takeaway: OpenAI gave the industry something rare: a vendor publishing its own detection failure curve next to the launch. Use it. Inventory your EU-facing text outputs, make an explicit dated decision on the Text provenance switch for each route, and log which model produced what. The mark is a signal for an expert holding a key, on long prose nobody edited. Everything else you need, you have to record yourself.
Primary sources
- Our approach to EU text provenance rules (OpenAI)
- textGrain: Entropy-Calibrated Watermarking for Language Model Text (OpenAI)
- Provenance signals in OpenAI-generated content (OpenAI)
- Article 50: Transparency obligations for providers and deployers of certain AI systems (EU AI Act Service Desk)
- A Watermark for Large Language Models (arXiv)
- How Claude marks AI-generated content (Anthropic)
- OpenAI explains how it will watermark ChatGPT (practitioner discussion) (Reddit)
- How AI text watermarking works (practitioner discussion) (Reddit)