OpenAI Unveils textGrain Invisible Text Watermarking: API Opt-In and Mandatory EU Rollout
OpenAI opened textGrain, a statistical watermarking algorithm that embeds imperceptible token distribution signatures into model outputs. API customers can opt in immediately, while ChatGPT and Codex accounts in the European Union will apply watermarks automatically to comply with EU AI Act provenance mandates.
OpenAI introduced textGrain on October 6, 2026, an imperceptible statistical watermarking technology designed to identify machine-generated text. Commercial API customers can now opt into textGrain across GPT-4o, GPT-6 Astra, and Codex endpoints.
Simultaneously, OpenAI confirmed that text generated in the European Union via ChatGPT and Codex will receive automatic, invisible textGrain watermarks in the coming weeks to satisfy disclosure obligations established by Article 50 of the EU Artificial Intelligence Act.
How textGrain Functions: Pseudorandom Green-List Perturbations
Traditional watermarking techniques—such as zero-width Unicode characters or rigid synonym substitutions—are fragile. They break as soon as a user copies text into a plain-text editor or runs a basic spellcheck.
textGrain operates at the sampling stage during auto-regressive generation:
┌────────────────────────────────────────────────────────────────────────┐
│ textGrain Statistical Watermarking Loop │
├────────────────────────────────────────────────────────────────────────┤
│ 1. Current Context: Sequence of preceding tokens (e.g., k=3 tokens) │
│ 2. Hash Seed: Compute HMAC-SHA256(context_window, secret_server_key) │
│ 3. Vocabulary Split: Seed partitions vocabulary into "Green" & "Red" │
│ 4. Logit Shift: Add small bias delta (+δ) to logits of Green tokens │
│ 5. Softmax & Sample: Model samples next token from biased distribution │
└────────────────────────────────────────────────────────────────────────┘
Because the bias delta is minimal ($delta approx 0.8$), the chosen token remains natural and contextually appropriate. However, an analysis of consecutive words reveals that the text selects "green-list" tokens with statistically impossible frequency ($p < 10^{-6}$), proving AI generation.
Verification and Robustness Metrics
OpenAI evaluated textGrain across academic essays, corporate press releases, and code scripts. The watermark survives substantial text manipulation:
| Text Modification Vector | Detection Accuracy (100 words) | Detection Accuracy (250 words) | False Positive Rate |
|---|---|---|---|
| Direct Unmodified Text | 94.8% | 99.9% | < 0.001% |
| Synonym Substitution (15% words) | 82.1% | 98.4% | < 0.001% |
| Light Human Paraphrasing | 76.5% | 94.2% | < 0.001% |
| Format Conversion (Markdown to Text) | 94.0% | 99.8% | < 0.001% |
Even when a student or writer changes adjectives or alters sentence order, the underlying distribution retains enough green-list tokens for statistical confirmation.
API Configuration and EU Compliance
Developers can enable textGrain by appending the watermark parameter to API calls:
import openai
client = openai.OpenAI()
response = client.chat.completions.create(
model="gpt-6-astra",
messages=[{"role": "user", "content": "Draft technical compliance summary."}],
watermark={"type": "textgrain", "strength": "standard"}
)
print(response.choices[0].message.content)
For EU end-users, this configuration will be injected automatically at the gateway layer, ensuring public documents, administrative drafts, and news releases generated by ChatGPT satisfy statutory transparency mandates.