Tools & Products

OpenAI Ultrafast Mode: 300+ Tokens/Sec Speculative Decoding Tier, API Benchmarks, and Latency Analysis

OpenAI revealed Ultrafast Mode at DevDay 2026, an API inference tier that accelerates GPT-6.1 Sol and Luna generation speeds up to 300+ tokens per second—a 14x improvement over standard endpoints. By combining multi-token prediction heads with specialized FP8 Blackwell kernels and speculative drafting models, Ultrafast Mode targets real-time autonomous agent loops and interactive voice systems.

By FreakVinci · 2026-09-29 · 16 min read

Inference Acceleration: Breaking the 300 Tokens/Second Barrier

At DevDay 2026, OpenAI unveiled Ultrafast Mode, an inference capability that clocks sustained output speeds between 280 and 320 tokens per second on frontier models. Designed to resolve the throughput bottlenecks that hamper multi-turn agent execution and interactive programming assistants, Ultrafast Mode accelerates model delivery by up to 14x compared to typical 20–25 tokens/second production endpoints.

Agentic systems often require 15 to 30 recursive model calls to complete a single code refactoring task or terminal debug session. At standard speeds, that loop can consume two to three minutes. Ultrafast Mode compresses total execution time down to less than 15 seconds.


Architectural Mechanisms Behind the Speedup

Achieving 300+ tokens per second on large transformer architectures requires overcoming memory bandwidth limits (the memory-bound nature of autoregressive generation). OpenAI engineered three synchronized infrastructure improvements:

Ultrafast Mode Inference Pipeline
┌───────────────────────────────────────────────────────────────┐
│ Client Request (Prompt Tokens)                                │
└──────────────────────┬────────────────────────────────────────┘
                       ▼
┌───────────────────────────────────────────────────────────────┐
│ Draft Engine (Distilled 3B Luna Draft Model)                 │
│ Proposes K=5 candidate tokens in single forward pass          │
└──────────────────────┬────────────────────────────────────────┘
                       ▼
┌───────────────────────────────────────────────────────────────┐
│ Target Model Verification (GPT-6.1 Sol on Blackwell B200)     │
│ Parallel verification of draft tree via FP8 Tensor Cores      │
│ Acceptance rate: 82% across coding and natural language       │
└──────────────────────┬────────────────────────────────────────┘
                       ▼
┌───────────────────────────────────────────────────────────────┐
│ Streamed Token Emission: 280-320 Tokens / Second              │
└───────────────────────────────────────────────────────────────┘
  1. Multi-Token Speculative Decoding: Instead of sampling one token per forward pass, a lightweight 3-billion-parameter draft model proposes blocks of 5 tokens. The primary GPT-6.1 Sol model evaluates all candidate tokens in a single parallel verification pass. Because verified tokens match the exact distribution of the full model, output quality remains numerically identical to standard sampling.
  2. Custom Blackwell FP8/FP4 Kernels: Running on NVIDIA GB200 NVL72 racks, the inference stack utilizes custom CUDA micro-kernels that keep KV-cache representations in high-bandwidth memory (HBM3e) without pipeline stalls.
  3. Speculative Tree-Search Caching: Common syntactical structures in languages like TypeScript, Python, and SQL achieve draft acceptance rates above 88%, resulting in theoretical peaks exceeding 350 tokens per second.

Empirical Latency and Throughput Benchmarks

Internal testing and independent developer benchmarks report the following latency deltas:

Model Configuration Output Speed (tok/s) Time-to-First-Token (TTFT) 1,000 Token Code Generation Time Cost Multiplier
GPT-6 Astra (Standard) 22 tok/s 340 ms 45.4 seconds 1.0x (Base)
Claude Opus 5.5 26 tok/s 310 ms 38.5 seconds 1.0x (Base)
Claude Sonnet 5.5 42 tok/s 210 ms 23.8 seconds 1.0x (Base)
GPT-6.1 Sol (Standard) 35 tok/s 180 ms 28.5 seconds 1.0x (Base)
GPT-6.1 Sol (Ultrafast) 308 tok/s 74 ms 3.2 seconds 1.4x (Premium)
GPT-6 Luna (Ultrafast) 345 tok/s 58 ms 2.9 seconds 1.3x (Premium)

Generating a 1,000-token modular React component or SQL migration drops from 28.5 seconds on standard Sol to 3.2 seconds on Ultrafast Mode.


Developer API Implementation

Enabling Ultrafast Mode in the OpenAI Node SDK requires specifying the speed tier flag:

import OpenAI from 'openai';

const openai = new OpenAI();

async function streamUltrafastAgentLoop() {
  const stream = await openai.chat.completions.create({
    model: 'gpt-6.1-sol',
    // @ts-ignore - OpenAI DevDay 2026 preview flag
    speed_tier: 'ultrafast',
    messages: [
      {
        role: 'system',
        content: 'You are an ultra-low-latency coding agent. Output complete executable scripts with zero conversational filler.'
      },
      {
        role: 'user',
        content: 'Write an asynchronous pipeline in Rust that reads from Kafka, validates protobuf schemas, and flushes to ClickHouse.'
      }
    ],
    stream: true,
    max_tokens: 4096
  });

  const startTime = performance.now();
  let tokenCount = 0;

  for await (const chunk of stream) {
    const text = chunk.choices[0]?.delta?.content || '';
    process.stdout.write(text);
    tokenCount++;
  }

  const durationSec = (performance.now() - startTime) / 1000;
  console.log(`\nGenerated ${tokenCount} tokens in ${durationSec.toFixed(2)}s (${(tokenCount / durationSec).toFixed(1)} tok/s)`);
}

Primary Target Applications

Ultrafast Mode unlocks distinct system workflows that were unfeasible at standard 25–40 tok/s rates:

  1. Sub-Second Agentic Compilers: Autonomous test runners that evaluate test failures, edit source files, and re-execute test runners in sub-10-second cycles.
  2. High-Frequency Financial Sentiment Classification: Parsing regulatory filings and earnings call transcripts across hundreds of live feeds concurrently.
  3. Conversational Speech Loops: Direct integration into real-time audio pipelines where synthesized speech text must stream ahead of phoneme audio generation to prevent buffer underruns.