OpenAI Leaks Hint at $500 Monthly ChatGPT Pro Max Plan: Fastest Work, Codex Acceleration, and Cerebras Infrastructure Ahead of DevDay 2026
Backend leaks inside ChatGPT web client code reveal an unannounced $500 monthly Pro Max subscription tier. Positioned above the paused $200 Pro plan, the tier introduces priority access to Fastest Work and Codex, 100GB workspace storage, an always-on assistant named o, and dedicated low-latency compute tied to Cerebras wafer-scale infrastructure.
OpenAI Leaks Hint at $500 Monthly ChatGPT Pro Max Plan
On September 24, 2026, reverse engineers analyzing OpenAI web client bundles uncovered references to an unannounced subscription tier named ChatGPT Pro Max. The configuration objects place the plan at $500 per month, with European tax-inclusive mockups reaching $600 per month.
The leak surfaces five days before OpenAI DevDay 2026, scheduled for September 29 in San Francisco. The discovery exposes a growing divide in compute allocation: as standard plans face rate limits and paused onboarding, OpenAI appears ready to monetize ultra-low-latency agent infrastructure for deep-pocketed enterprise engineers and quantitative trading desks.
1. Discovery and Leaked Frontend Artifacts
The initial discovery originated from inspections of minified JavaScript chunks serving the ChatGPT account settings view. Within the plan selection modal schema, a new entry named chatgpt_pro_max_monthly sat directly above the existing chatgpt_pro_monthly tier.
{
"plan_id": "chatgpt_pro_max_monthly",
"display_name": "ChatGPT Pro Max",
"pricing": {
"base_usd": 500,
"interval": "month",
"tax_handling": "exclusive",
"regional_ceilings": {
"GB": 490,
"EU": 580,
"DE": 595
}
},
"entitlements": {
"compute_lane": "fastest_dedicated",
"work_agent_concurrency": 8,
"codex_agent_concurrency": 16,
"memory_context_mode": "maximum_persistent",
"workspace_storage_gb": 100,
"always_on_agent_enabled": true,
"flag_assistant_o": true
}
}
The bundle labels the tier's primary selling proposition as "Fastest Work and Codex." Additional copy strings retrieved from localization tables include:
- "Uncapped generation velocity for multi-hour research and development workloads."
- "Direct dispatch to specialized hardware fabrics with zero queuing delay."
- "Includes o, your always-on assistant for autonomous inbox and workspace actions."
2. Subscription Hierarchy and Pricing Economics
The leaked $500 tier reorganizes OpenAI's pricing structure into a steep four-tier pyramid:
| Plan Tier | Monthly Price | Target User | Dedicated Compute Lane | Agent Concurrency | Quota Architecture |
|---|---|---|---|---|---|
| Free | $0 | Casual Queries | Standard Shared Fleet | None | Dynamic throttling during peak hours |
| Plus | $20 | Knowledge Workers | Shared Priority Fleet | 1 Active Agent | 40-80 messages / 3 hours on flagship models |
| Pro | $200 | Power Developers & Researchers | Priority Compute Pool | 4 Active Agents | Unlimited standard queries; high reasoning limits |
| Pro Max (Leaked) | $500 ($600 w/ VAT) | Quantitative Engineers & Autonomous Labs | Dedicated Wafer-Scale Fabric | 16 Active Agents | Uncapped generation; lowest latency guarantees |
The $500 monthly figure represents a 150% jump over the $200 Pro plan and a 25-fold multiplier over the $20 Plus subscription. In European markets with standard 20% to 22% value-added taxes, monthly invoices will land between $580 and $600, or roughly $7,200 annually.
This pricing reflects real server operational expenses. Running multi-turn reasoning loops with deep tree search consumes hundreds of thousands of hidden thinking tokens per prompt. When an autonomous coding agent edits dozens of files, runs unit tests, parses tracebacks, and iterates across hours, the computational load equals thousands of standard conversational interactions.
3. Dissecting "Fastest Work and Codex"
The leaked entitlement string specifies "Fastest Work and Codex." To understand why this requires a dedicated pricing tier, one must look at how OpenAI categorizes agent operations.
Work: Autonomous Long-Horizon Planning
In OpenAI's internal product taxonomy, "Work" represents autonomous background execution. Unlike standard chat completions where the user waits for a single response, Work tasks run for 10 minutes to four hours.
An agent assigned to "Conduct a competitor patent audit on solid-state battery anodes" generates a plan, queries external databases, parses PDF filings, compares claims, builds vector indexes in memory, and synthesizes a 40-page report. This requires persistent memory and hundreds of sequential tool calls.
Codex: Repository-Wide Mutation
Codex has evolved from a simple code-completion model into a full-context engineering agent. In its current implementation, Codex operates across an entire workspace:
- Clones a target repository into an isolated sandboxed container.
- Ingests file dependency graphs and project configuration files.
- Generates regression test matrices.
- Executes builds, captures exit codes, and self-corrects broken assertions.
In standard shared clusters, running 16 concurrent Codex agents creates severe queue contention. The Pro Max tier reserves dedicated execution lanes, ensuring agent steps complete in seconds rather than stalling behind public batch requests.
4. Hardware Fabric: The Cerebras Wafer-Scale Connection
The term "Fastest" in the leaked telemetry strings points directly to specialized silicon. Industry analysts and supply chain reports indicate OpenAI runs its lowest-latency inference pipelines on hardware manufactured by Cerebras Systems.
Traditional GPU clusters built on NVIDIA H100 and Blackwell B200 systems depend on High Bandwidth Memory (HBM3e). While GPUs deliver high parallel floating-point throughput, autoregressive token generation remains memory-bandwidth bound at small batch sizes. For an interactive agent generating code step by step, standard GPUs hit token throughput walls between 50 and 90 tokens per second.
Token Generation Throughput Comparison (Tokens Per Second):
--------------------------------------------------------------------------------
Standard GPU Cloud Fleet (H100/H200, Shared Batching): [65 tok/s]
Priority GPU Dedicated Lane (Pro Tier, Isolated Node): [140 tok/s]
Cerebras Wafer-Scale Engine (CS-3, On-Chip SRAM): [1,150 tok/s]
--------------------------------------------------------------------------------
The Cerebras CS-3 Wafer-Scale Engine solves memory bottlenecks by fabricating an entire 300mm silicon wafer as a single contiguous chip:
- 900,000 AI cores printed on a single piece of silicon.
- 44 gigabytes of on-chip SRAM, avoiding off-chip memory busses entirely.
- 1.2 petabytes per second of internal memory bandwidth.
OpenAI previously deployed Cerebras clusters for Codex-Spark and GPT-5.6 Sol Ultrafast, sustaining generation rates between 750 and 1,200 tokens per second. In autonomous agent loops, a 10x increase in generation speed turns a 20-minute test refactor into a 90-second run. For elite engineering organizations, saving 18 minutes on every commit justifies the $500 monthly fee.
5. Leaked Always-On Assistant: Codename "o"
Beyond speed improvements, the frontend code exposes feature flags for an autonomous companion labeled "o, your always-on assistant."
interface AssistantOProfile {
identity_slug: "assistant_o";
mode: "persistent_daemon";
assigned_inbox: string; // e.g., user.agent@chatgpt.work
permissions: {
calendar_write: boolean;
workspace_fs_sync: boolean;
webhook_listener: boolean;
background_polling_minutes: number; // default: 5
};
state_persistence: "infinite_ttl";
}
Unlike standard ChatGPT sessions that sleep when the browser tab closes, Assistant "o" operates continuously in cloud containers:
- Assigned Mailbox: Each Pro Max user receives a dedicated
@chatgpt.workforwarding address. Inbound emails from clients or monitoring systems automatically trigger agent reasoning routines. - Continuous Environment Monitoring: The agent monitors GitHub webhooks, server error logs, and calendar schedules, preparing action summaries before the user sits at their desk.
- Proactive Interventions: If a continuous integration pipeline fails at 3:00 AM, the agent pulls the error logs, writes a fix, verifies it against unit tests, and presents an approval card to the developer.
This shifts the interaction model from user-initiated query-response to continuous autonomous collaboration.
6. Developer Reactions and the "Luxury AI" Divide
The leak arrived during widespread user frustration over OpenAI's subscription infrastructure.
The September 10 Sign-Up Freeze
On September 10, 2026, OpenAI abruptly suspended new registrations for the $200 monthly Pro plan. Visitors attempting to upgrade received a waitlist notice citing cluster capacity constraints:
"Due to overwhelming compute demand following the launch of GPT-6 Astra, new ChatGPT Pro subscriptions are temporarily paused. Existing subscribers retain uninterrupted access."
Existing Pro subscribers reported recurring delays in service quota resets. Instead of instantaneous usage replenishments at the start of each billing month, background jobs lagged by up to 36 hours, leaving paying developers throttled on heavy reasoning workloads.
Community Friction Over Tier Gating
News of a $500 plan ignited intense debate across developer forums:
- The Cost Polarization: A developer spending $20 per month receives throttled access during peak hours. The next functional upgrade was $200, and is now closed. Introducing a $500 plan suggests frontier models are shifting into corporate luxury goods.
- Compute Realism: Enterprise founders countered that renting eight dedicated H100 instances on cloud providers costs $15,000 to $25,000 per month. A turnkey $500 service that handles infrastructure, wafer-scale routing, and agent coordination represents massive savings for funded teams.
- The Open-Source Alternative: Platforms such as OpenCode offering $60 monthly allowances on open-weight models like DeepSeek V4.1 Flash provide a viable escape hatch for independent engineers resisting closed ecosystem subscription inflation.
7. Plan Feature Comparison Matrix
The table below contrasts the leaked specifications of ChatGPT Pro Max against current production tiers:
| Technical Feature | ChatGPT Plus ($20/mo) | ChatGPT Pro ($200/mo) | ChatGPT Pro Max ($500/mo, Leaked) |
|---|---|---|---|
| Primary Compute Backbone | Shared Azure GPU Fleet | Priority GPU Allocations | Dedicated Cerebras / GPU Low-Latency Mesh |
| Output Token Velocity | 40 - 70 tokens/sec | 80 - 150 tokens/sec | 800 - 1,200+ tokens/sec |
| Agent Workspace Storage | 2 GB shared | 25 GB persistent | 100 GB dedicated high-IOPS NVMe |
| Max Context Window | 128,000 tokens | 200,000 tokens | Up to 1,000,000 tokens with full KV caching |
| Autonomous Daemon (o) | Not Supported | Experimental Preview | Full Background Runtime & Dedicated Email |
| Codex Concurrency | 1 process | 4 processes | 16 processes |
| Availability Status | Open | Paused / Waitlist | Planned DevDay Announcement |
8. What to Watch for at DevDay 2026
OpenAI DevDay 2026 opens Monday, September 29, at the Bill Graham Civic Auditorium in San Francisco. CEO Sam Altman will deliver the opening keynote at 10:00 AM Pacific Time.
Key technical announcements expected during the keynote include:
- Official Launch of Pro Max: Confirmation of the $500 price point, availability dates, and rollout strategy for users stuck on the Pro waitlist.
- Cerebras Partnership Expansion: Public architectural disclosures detailing how wafer-scale silicon integrates into OpenAI's hybrid Azure routing mesh.
- General Availability of Assistant "o": Live demonstrations of background autonomous agents managing email triage, Jira sprint updates, and Git conflict resolution.
- Agent SDK & Workspace APIs: New API primitives allowing third-party developers to embed long-horizon Work agents into native operating system applications.
The outcome of DevDay will determine whether OpenAI can satisfy compute demand while maintaining trust among independent software developers who built their businesses on accessible AI infrastructure.