Perplexity Decisions API: Deep Technical Guide to pplx-decider-v1-27b, Sub-50ms Routing, and Micro-Cent Pricing
Perplexity launched its Decisions API on October 1, 2026, offering direct programmatic access to pplx-decider-v1-27b. Operating at $0.10 per million decisions with sub-50ms response times, the API outputs normalized probability distributions across user-defined candidate classes, replacing slow autoregressive LLM classifiers in production RAG systems.
Overview: Why Perplexity Built a Dedicated Decisions Endpoint
When building high-throughput search engines, multi-hop retrieval-augmented generation (RAG) systems, or agent coordination networks, engineering teams constantly make categorical choices:
- Does this user query require live internet crawling or local database lookup?
- Is this customer inquiry a technical bug, a billing dispute, or sales spam?
- Should this trading bot buy, hold, or liquidate?
Using a generative model like gpt-4o or claude-3-5-sonnet to answer a multiple-choice question introduces 600ms to 1,500ms of latency and costs dollars per day under production load.
On October 1, 2026, Perplexity AI opened public developer access to the Decisions API (https://api.perplexity.ai/v1/decisions), powered by pplx-decider-v1-27b. The endpoint executes decisions in under 50 milliseconds at a fixed rate of $0.10 per million decisions.
Perplexity Decisions API Topology
┌────────────────────────────────────────────────────────┐
│ Client Microservice (API Gateway / Search Dispatcher) │
└──────────────────────────┬─────────────────────────────┘
│ POST /v1/decisions (Input + Candidates)
▼
┌────────────────────────────────────────────────────────┐
│ Perplexity pplx-decider-v1-27b Headless Inference │
│ Single Forward Pass • Zero Token Decoding Latency │
└──────────────────────────┬─────────────────────────────┘
│ 32 ms Round-Trip Time
▼
┌────────────────────────────────────────────────────────┐
│ Structured JSON: { decision: "live_search", p: 0.984 } │
└────────────────────────────────────────────────────────┘
API Specification and Schema
Request Headers
Authorization: Bearer <PERPLEXITY_API_KEY>Content-Type: application/json
Request Payload
{
"model": "pplx-decider-v1-27b",
"input": "User asks: Can you help me calculate the internal rate of return on my real estate investment spreadsheet?",
"candidates": [
"financial_calculator_tool",
"web_search_crawler",
"conversational_explainer",
"security_refusal_filter"
],
"temperature": 0.0
}
Response Payload
{
"id": "dec_8f7b2c91a0e4",
"object": "decision",
"created": 1790892400,
"model": "pplx-decider-v1-27b",
"decision": "financial_calculator_tool",
"confidence": 0.9842,
"distribution": {
"financial_calculator_tool": 0.9842,
"conversational_explainer": 0.0121,
"web_search_crawler": 0.0034,
"security_refusal_filter": 0.0003
},
"latency_ms": 31
}
Python Production Implementation
The following script integrates the Decisions API into a production RAG routing pipeline:
import requests
import os
import time
PERPLEXITY_API_KEY = os.environ.get("PERPLEXITY_API_KEY")
def route_user_query(query: str) -> str:
url = "https://api.perplexity.ai/v1/decisions"
headers = {
"Authorization": f"Bearer {PERPLEXITY_API_KEY}",
"Content-Type": "application/json"
}
payload = {
"model": "pplx-decider-v1-27b",
"input": query,
"candidates": [
"vector_kb_search",
"live_google_news",
"code_interpreter_sandbox",
"direct_llm_answer"
]
}
start = time.perf_counter()
response = requests.post(url, json=payload, headers=headers, timeout=1.0)
latency = (time.perf_counter() - start) * 1000
if response.status_code == 200:
data = response.json()
print(f"Decided: {data['decision']} (p={data['confidence']:.3f}) in {latency:.1f}ms")
return data["decision"]
else:
raise RuntimeError(f"Decisions API failed: {response.text}")
# Test execution
route = route_user_query("What happened in the US Senate vote on AI safety regulations this afternoon?")
TypeScript Implementation for Next.js and Cloudflare Workers
export async function getDecision(query: string, candidates: string[]) {
const res = await fetch('https://api.perplexity.ai/v1/decisions', {
method: 'POST',
headers: {
'Authorization': `Bearer ${process.env.PERPLEXITY_API_KEY}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({
model: 'pplx-decider-v1-27b',
input: query,
candidates
})
});
if (!res.ok) {
throw new Error(`Perplexity API error: ${await res.text()}`);
}
const result = await res.json();
return {
decision: result.decision as string,
confidence: result.confidence as number,
latencyMs: result.latency_ms as number
};
}
Economics: Decisions API vs Standard Generative Endpoints
For a consumer platform processing 10,000,000 user routing requests per month:
| Model / Endpoint | Latency (P50) | Monthly Cost (10M requests) | Annual Expenditure |
|---|---|---|---|
| Perplexity Decisions API | 31 ms | $1.00 | $12.00 / yr |
| OpenAI GPT-4o-mini (JSON) | 480 ms | $65.00 | $780.00 / yr |
| Claude 3.5 Haiku | 520 ms | $105.00 | $1,260.00 / yr |
| Claude 3.7 Sonnet | 890 ms | $750.00 | $9,000.00 / yr |
The Decisions API lowers routing expenditure by 98.4% while executing 15x faster than lightweight LLMs, making it the most cost-effective decision engine for production microservice architectures.