Tools & Products

Perplexity Decisions API: Deep Technical Guide to pplx-decider-v1-27b, Sub-50ms Routing, and Micro-Cent Pricing

Perplexity launched its Decisions API on October 1, 2026, offering direct programmatic access to pplx-decider-v1-27b. Operating at $0.10 per million decisions with sub-50ms response times, the API outputs normalized probability distributions across user-defined candidate classes, replacing slow autoregressive LLM classifiers in production RAG systems.

By FreakVinci · 2026-10-01 · 13 min read

Overview: Why Perplexity Built a Dedicated Decisions Endpoint

When building high-throughput search engines, multi-hop retrieval-augmented generation (RAG) systems, or agent coordination networks, engineering teams constantly make categorical choices:

  • Does this user query require live internet crawling or local database lookup?
  • Is this customer inquiry a technical bug, a billing dispute, or sales spam?
  • Should this trading bot buy, hold, or liquidate?

Using a generative model like gpt-4o or claude-3-5-sonnet to answer a multiple-choice question introduces 600ms to 1,500ms of latency and costs dollars per day under production load.

On October 1, 2026, Perplexity AI opened public developer access to the Decisions API (https://api.perplexity.ai/v1/decisions), powered by pplx-decider-v1-27b. The endpoint executes decisions in under 50 milliseconds at a fixed rate of $0.10 per million decisions.

Perplexity Decisions API Topology
┌────────────────────────────────────────────────────────┐
│ Client Microservice (API Gateway / Search Dispatcher)  │
└──────────────────────────┬─────────────────────────────┘
                           │ POST /v1/decisions (Input + Candidates)
                           ▼
┌────────────────────────────────────────────────────────┐
│ Perplexity pplx-decider-v1-27b Headless Inference      │
│ Single Forward Pass • Zero Token Decoding Latency      │
└──────────────────────────┬─────────────────────────────┘
                           │ 32 ms Round-Trip Time
                           ▼
┌────────────────────────────────────────────────────────┐
│ Structured JSON: { decision: "live_search", p: 0.984 } │
└────────────────────────────────────────────────────────┘

API Specification and Schema

Request Headers

  • Authorization: Bearer <PERPLEXITY_API_KEY>
  • Content-Type: application/json

Request Payload

{
  "model": "pplx-decider-v1-27b",
  "input": "User asks: Can you help me calculate the internal rate of return on my real estate investment spreadsheet?",
  "candidates": [
    "financial_calculator_tool",
    "web_search_crawler",
    "conversational_explainer",
    "security_refusal_filter"
  ],
  "temperature": 0.0
}

Response Payload

{
  "id": "dec_8f7b2c91a0e4",
  "object": "decision",
  "created": 1790892400,
  "model": "pplx-decider-v1-27b",
  "decision": "financial_calculator_tool",
  "confidence": 0.9842,
  "distribution": {
    "financial_calculator_tool": 0.9842,
    "conversational_explainer": 0.0121,
    "web_search_crawler": 0.0034,
    "security_refusal_filter": 0.0003
  },
  "latency_ms": 31
}

Python Production Implementation

The following script integrates the Decisions API into a production RAG routing pipeline:

import requests
import os
import time

PERPLEXITY_API_KEY = os.environ.get("PERPLEXITY_API_KEY")

def route_user_query(query: str) -> str:
    url = "https://api.perplexity.ai/v1/decisions"
    headers = {
        "Authorization": f"Bearer {PERPLEXITY_API_KEY}",
        "Content-Type": "application/json"
    }
    payload = {
        "model": "pplx-decider-v1-27b",
        "input": query,
        "candidates": [
            "vector_kb_search",
            "live_google_news",
            "code_interpreter_sandbox",
            "direct_llm_answer"
        ]
    }
    
    start = time.perf_counter()
    response = requests.post(url, json=payload, headers=headers, timeout=1.0)
    latency = (time.perf_counter() - start) * 1000
    
    if response.status_code == 200:
        data = response.json()
        print(f"Decided: {data['decision']} (p={data['confidence']:.3f}) in {latency:.1f}ms")
        return data["decision"]
    else:
        raise RuntimeError(f"Decisions API failed: {response.text}")

# Test execution
route = route_user_query("What happened in the US Senate vote on AI safety regulations this afternoon?")

TypeScript Implementation for Next.js and Cloudflare Workers

export async function getDecision(query: string, candidates: string[]) {
  const res = await fetch('https://api.perplexity.ai/v1/decisions', {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${process.env.PERPLEXITY_API_KEY}`,
      'Content-Type': 'application/json'
    },
    body: JSON.stringify({
      model: 'pplx-decider-v1-27b',
      input: query,
      candidates
    })
  });

  if (!res.ok) {
    throw new Error(`Perplexity API error: ${await res.text()}`);
  }

  const result = await res.json();
  return {
    decision: result.decision as string,
    confidence: result.confidence as number,
    latencyMs: result.latency_ms as number
  };
}

Economics: Decisions API vs Standard Generative Endpoints

For a consumer platform processing 10,000,000 user routing requests per month:

Model / Endpoint Latency (P50) Monthly Cost (10M requests) Annual Expenditure
Perplexity Decisions API 31 ms $1.00 $12.00 / yr
OpenAI GPT-4o-mini (JSON) 480 ms $65.00 $780.00 / yr
Claude 3.5 Haiku 520 ms $105.00 $1,260.00 / yr
Claude 3.7 Sonnet 890 ms $750.00 $9,000.00 / yr

The Decisions API lowers routing expenditure by 98.4% while executing 15x faster than lightweight LLMs, making it the most cost-effective decision engine for production microservice architectures.