OpenAI Decisions API: Sub-Second Classification, Model Routing, and High-Throughput Inference Architecture
OpenAI unveiled the Decisions API at DevDay 2026, delivering sub-25ms classification and semantic routing. Powered by distilled Luna neural heads that bypass standard autoregressive token generation, the API enables high-throughput request dispatching, real-time safety gating, and programmatic ad decisioning in production infrastructure.
Sub-Second Routing: Solving the LLM Ingress Bottleneck
At DevDay 2026, OpenAI introduced the Decisions API, a specialized endpoint designed to classify, tag, and route incoming requests in under 25 milliseconds. Traditional large language models incur substantial latency penalties (often 200–500ms) when asked simple routing questions such as "Is this query a programming bug or a billing question?"
By decoupling classification from autoregressive text generation, the Decisions API allows engineering teams to deploy intelligent routing logic at API gateways without degrading user perceived response times.
Non-Autoregressive Classifier Architecture
The Decisions API operates differently from conventional chat completion endpoints:
Decisions API Execution Architecture
┌─────────────────────────────────────────────────────────────┐
│ Ingress Query (User Prompt / Payload) │
└──────────────────────┬──────────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────┐
│ Distilled Luna 1.8B Encoder Pass │
│ Single parallel forward pass across input embeddings │
└──────────────────────┬──────────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────┐
│ Categorical Softmax Projection Layer │
│ Computes log-probabilities across defined decision schemas │
│ Latency: 18 - 24 ms total inference time │
└──────────────────────┬──────────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────┐
│ Typed Decision Output & Route Dispatch │
│ - Category: 'code_generation' (Confidence: 0.982) │
│ - Recommended Route: 'gpt-6.1-sol' │
└─────────────────────────────────────────────────────────────┘
- Non-Autoregressive Projection: Rather than sampling sequential tokens, input text passes through a frozen transformer backbone paired with an output linear probe that emits class probabilities directly.
- Deterministic Schema Binding: Developers define categorical enums directly in the request payload. The model computes normalized logits strictly over the provided enum values.
- Edge Deployment: OpenAI serves the Decisions API from regional edge points-of-presence (PoPs), terminating TLS connections close to user infrastructure.
Latency and Throughput Comparison
Benchmarking against traditional LLM classification pipelines reveals the architectural difference:
| Method | Backbone Model | Median Latency (P50) | Tail Latency (P99) | Cost per 100K Decisions |
|---|---|---|---|---|
| Chat Completion JSON | GPT-6.1 Sol | 340 ms | 680 ms | $1.25 |
| Structured Output | GPT-6 Luna | 120 ms | 240 ms | $0.20 |
| Custom FastText Container | Self-Hosted CPU | 8 ms | 45 ms | Infrastructure fixed |
| OpenAI Decisions API | Distilled Luna Pro | 19 ms | 34 ms | $0.10 |
The Decisions API delivers 18x lower latency than standard structured outputs on Sol while eliminating the engineering overhead of training and deploying bespoke self-hosted classifier microservices.
Implementation Example: Dynamic Gateway Router
A typical Node.js reverse proxy implementation illustrates how developers use the Decisions API to dispatch traffic:
import OpenAI from 'openai';
const openai = new OpenAI();
interface RoutingDecision {
destination_model: 'gpt-6.1-sol' | 'gpt-6-luna' | 'local_cache' | 'block_policy';
confidence: number;
}
export async function routeUserPrompt(userPrompt: string): Promise<string> {
// @ts-ignore - OpenAI DevDay 2026 preview endpoint
const decision = await openai.decisions.create({
input: userPrompt,
schema: {
type: 'categorical',
categories: [
'complex_coding_or_math',
'simple_conversational',
'cached_faq',
'policy_violation'
]
}
});
const category = decision.selected_category;
switch (category) {
case 'complex_coding_or_math':
return 'gpt-6.1-sol';
case 'simple_conversational':
return 'gpt-6-luna';
case 'cached_faq':
return 'local_cache';
default:
return 'block_policy';
}
}
Programmatic Ad Matching and Enterprise Integrations
In addition to API routing, OpenAI announced integrations with advertising infrastructure providers like Kevel. In programmatic ad decisioning, ad servers have less than 50 milliseconds to parse user query context, match target advertiser intents, and submit bids. The Decisions API operates within this tight RTB (Real-Time Bidding) budget, enabling contextual keyword matching without tracking cookies.