Tools & Products

OpenAI Decisions API: Sub-Second Classification, Model Routing, and High-Throughput Inference Architecture

OpenAI unveiled the Decisions API at DevDay 2026, delivering sub-25ms classification and semantic routing. Powered by distilled Luna neural heads that bypass standard autoregressive token generation, the API enables high-throughput request dispatching, real-time safety gating, and programmatic ad decisioning in production infrastructure.

By FreakVinci · 2026-09-29 · 15 min read

Sub-Second Routing: Solving the LLM Ingress Bottleneck

At DevDay 2026, OpenAI introduced the Decisions API, a specialized endpoint designed to classify, tag, and route incoming requests in under 25 milliseconds. Traditional large language models incur substantial latency penalties (often 200–500ms) when asked simple routing questions such as "Is this query a programming bug or a billing question?"

By decoupling classification from autoregressive text generation, the Decisions API allows engineering teams to deploy intelligent routing logic at API gateways without degrading user perceived response times.


Non-Autoregressive Classifier Architecture

The Decisions API operates differently from conventional chat completion endpoints:

Decisions API Execution Architecture
┌─────────────────────────────────────────────────────────────┐
│ Ingress Query (User Prompt / Payload)                       │
└──────────────────────┬──────────────────────────────────────┘
                       ▼
┌─────────────────────────────────────────────────────────────┐
│ Distilled Luna 1.8B Encoder Pass                            │
│ Single parallel forward pass across input embeddings        │
└──────────────────────┬──────────────────────────────────────┘
                       ▼
┌─────────────────────────────────────────────────────────────┐
│ Categorical Softmax Projection Layer                        │
│ Computes log-probabilities across defined decision schemas  │
│ Latency: 18 - 24 ms total inference time                    │
└──────────────────────┬──────────────────────────────────────┘
                       ▼
┌─────────────────────────────────────────────────────────────┐
│ Typed Decision Output & Route Dispatch                      │
│ - Category: 'code_generation' (Confidence: 0.982)           │
│ - Recommended Route: 'gpt-6.1-sol'                          │
└─────────────────────────────────────────────────────────────┘
  1. Non-Autoregressive Projection: Rather than sampling sequential tokens, input text passes through a frozen transformer backbone paired with an output linear probe that emits class probabilities directly.
  2. Deterministic Schema Binding: Developers define categorical enums directly in the request payload. The model computes normalized logits strictly over the provided enum values.
  3. Edge Deployment: OpenAI serves the Decisions API from regional edge points-of-presence (PoPs), terminating TLS connections close to user infrastructure.

Latency and Throughput Comparison

Benchmarking against traditional LLM classification pipelines reveals the architectural difference:

Method Backbone Model Median Latency (P50) Tail Latency (P99) Cost per 100K Decisions
Chat Completion JSON GPT-6.1 Sol 340 ms 680 ms $1.25
Structured Output GPT-6 Luna 120 ms 240 ms $0.20
Custom FastText Container Self-Hosted CPU 8 ms 45 ms Infrastructure fixed
OpenAI Decisions API Distilled Luna Pro 19 ms 34 ms $0.10

The Decisions API delivers 18x lower latency than standard structured outputs on Sol while eliminating the engineering overhead of training and deploying bespoke self-hosted classifier microservices.


Implementation Example: Dynamic Gateway Router

A typical Node.js reverse proxy implementation illustrates how developers use the Decisions API to dispatch traffic:

import OpenAI from 'openai';

const openai = new OpenAI();

interface RoutingDecision {
  destination_model: 'gpt-6.1-sol' | 'gpt-6-luna' | 'local_cache' | 'block_policy';
  confidence: number;
}

export async function routeUserPrompt(userPrompt: string): Promise<string> {
  // @ts-ignore - OpenAI DevDay 2026 preview endpoint
  const decision = await openai.decisions.create({
    input: userPrompt,
    schema: {
      type: 'categorical',
      categories: [
        'complex_coding_or_math',
        'simple_conversational',
        'cached_faq',
        'policy_violation'
      ]
    }
  });

  const category = decision.selected_category;

  switch (category) {
    case 'complex_coding_or_math':
      return 'gpt-6.1-sol';
    case 'simple_conversational':
      return 'gpt-6-luna';
    case 'cached_faq':
      return 'local_cache';
    default:
      return 'block_policy';
  }
}

Programmatic Ad Matching and Enterprise Integrations

In addition to API routing, OpenAI announced integrations with advertising infrastructure providers like Kevel. In programmatic ad decisioning, ad servers have less than 50 milliseconds to parse user query context, match target advertiser intents, and submit bids. The Decisions API operates within this tight RTB (Real-Time Bidding) budget, enabling contextual keyword matching without tracking cookies.