OpenAI Boosts GPT-6 Astra and GPT-6.1 Sol Speed by 50% in ChatGPT, Launching 28-Day Daily Upgrade Cycle
OpenAI deployed hardware-level inference optimizations that cut response latency in half for GPT-6 Astra and GPT-6.1 Sol. The rollout initiates a 28-day continuous improvement program delivering daily engine refinements to ChatGPT Codex and Work enterprise subscribers.
OpenAI pushed an inference acceleration update to GPT-6 Astra and GPT-6.1 Sol inside ChatGPT on October 6, 2026. The update cuts generation latency by approximately 50%, doubling output throughput for conversational answers and multi-file code refactoring.
Alongside the speed increase, OpenAI announced a 28-Day Daily Improvement Pledge. For the next four weeks, the engineering team will deploy daily production updates targeting inference stability, context cache efficiency, and agentic tool invocation across the Codex and ChatGPT Work tiers.
Benchmark Measurements: Throughput and Time-to-First-Token
Telemetry gathered across 50,000 production prompts demonstrates significant performance improvements over the September baseline:
| Metric | GPT-6 Astra (Sept 2026) | GPT-6 Astra (Oct 6, 2026) | GPT-6.1 Sol (Sept 2026) | GPT-6.1 Sol (Oct 6, 2026) |
|---|---|---|---|---|
| Output Speed (Tokens/Sec) | 62 tps | 124 tps (+100%) | 95 tps | 188 tps (+97.8%) |
| Time-to-First-Token (TTFT) | 910 ms | 440 ms (-51.6%) | 620 ms | 310 ms (-50.0%) |
| SWE-bench Execution Latency | 14.2 min | 7.1 min (-50.0%) | 9.8 min | 4.9 min (-50.0%) |
| Context Prefill (128K context) | 3.4 sec | 1.2 sec (-64.7%) | 2.1 sec | 0.8 sec (-61.9%) |
The halving of end-to-end execution time directly impacts autonomous agent loops. Complex coding workflows that previously required fifteen minutes of waiting now conclude in under seven minutes.
Technical Mechanism: Speculative Kernels on GB200 NVL72
The latency drop stems from three coordinated systems modifications:
- Speculative Draft Verification: OpenAI paired GPT-6 frontier checkpoints with an in-memory 7B draft model. The draft model proposes sequences of 4 to 6 candidate tokens, which the primary 1T+ model validates in a single parallel tensor pass.
- Cluster Migration to Blackwell GB200: Workloads moved from Hopper H100 pods to newly integrated NVIDIA GB200 NVL72 racks, leveraging 1.8 TB/s NVLink bi-directional interconnect bandwidth.
- Prefix Caching v3: Repository file trees, system prompts, and MCP tool declarations stay resident in GPU high-bandwidth memory (HBM3e) across repeated user turns.
┌────────────────────────────────────────────────────────────────────────┐
│ Speculative Decoding Acceleration Pipeline │
├────────────────────────────────────────────────────────────────────────┤
│ Prompt Ingest ──► Fast 7B Draft Engine (Generates 5 tokens in 12ms) │
│ │ │
│ ▼ │
│ GPT-6 Frontier Primary Engine │
│ Validates 5 candidate tokens in 1 GPU step (18ms) │
│ │ │
│ ▼ │
│ Acceptance Rate: 84% ──► Effective Speed: 188 tokens/sec │
└────────────────────────────────────────────────────────────────────────┘
The 28-Day Daily Upgrade Schedule
The 28-day campaign commits OpenAI to daily, non-disruptive production changes through November 3, 2026. Key targets announced for Codex and Work users include:
- Days 1–7: Terminal output streaming latency and shell command execution verification.
- Days 8–14: Memory footprint reductions for multi-repository context indexing.
- Days 15–21: Sub-second tool-calling handoffs for MCP and custom actions.
- Days 22–28: Extended context reasoning stability up to 1 million active tokens.