Research

Cloudflare Drops Clef and Clef-flash: Open Fast Decision Models Returning Calibrated Probabilities Over Text

Cloudflare released open weights for Clef and Clef-flash, a family of non-autoregressive decision models trained specifically to return calibrated floating-point probability distributions rather than generative text strings in under 4 milliseconds on edge Workers.

By Julian Thorne · 2026-10-03 · 13 min read

Cloudflare open-sourced Clef and Clef-flash, two purpose-built decision models designed to solve one of the most wasteful anti-patterns in modern software engineering: prompting massive autoregressive language models just to receive a binary decision or category label.

Instead of generating tokenized text strings (such as {"action": "route_to_sql", "confidence": 0.92}), Clef models output calibrated floating-point probability distributions directly across discrete choice tensors in a single forward pass.


Latency and Efficiency Comparison

Using an autoregressive LLM for system classification incurs token decoding latency, prompt serialization overhead, and JSON parsing failures. Clef eliminates the token generation loop entirely:

Model Architecture Parameter Size Mean Inference Time Output Format Expected Calibration Error (ECE)
GPT-4o Mini (JSON Mode) Undisclosed 220 ms - 450 ms JSON text string 0.084 (Overconfident)
Llama-3.1-8B-Instruct 8.0 Billion 115 ms - 280 ms Structured text 0.062
Clef (Base Model) 350 Million 4.2 ms Normalized Float32 Tensor 0.014 (Highly Calibrated)
Clef-flash 95 Million 1.8 ms Normalized Float32 Tensor 0.018 (Highly Calibrated)

Mathematical Mechanism: Direct Decision Projections

Autoregressive models compute probabilities over a 128,000-token vocabulary sequentially. Clef, by contrast, maps input representations directly into a low-dimensional decision simplex:

[Input Prompt / Request Payload]
               │
               ▼
┌─────────────────────────────────────────┐
│ Transformer Encoder Backbone            │
│ (Bidirectional Attention, 12 Layers)    │
└──────────────────────┬──────────────────┘
                       │
                       ▼ (Pooled Latent Representation)
┌─────────────────────────────────────────┐
│ Dense Projection Calibration Head       │
│ • Temperature-scaled Softmax            │
│ • Mathematical bounds check             │
└──────────────────────┬──────────────────┘
                       │
                       ▼
        [ [0.942, 0.041, 0.017] ]
   (Probability Distribution Over Branches)

Because Clef requires only a single encoder pass without iterative auto-regressive unrolling, memory bandwidth saturation drops to near zero.


Deploying Clef-flash in Cloudflare Workers

Developers can deploy Clef directly inside a Cloudflare Worker using native Workers AI bindings:

export interface Env {
  AI: any;
}

export default {
  async fetch(request: Request, env: Env): Promise<Response> {
    const { userQuery } = await request.json();

    // Evaluate routing decision in sub-2ms
    const decision = await env.AI.run('@cf/cloudflare/clef-flash', {
      input: userQuery,
      candidates: ['retrieve_rag_docs', 'execute_code_sandbox', 'direct_answer']
    });

    // Returns exact calibrated probabilities
    // decision = { probabilities: [0.892, 0.085, 0.023], predictedIndex: 0 }
    if (decision.probabilities[0] > 0.85) {
      return Response.json({ route: 'rag', confidence: decision.probabilities[0] });
    }

    return Response.json({ route: 'direct', confidence: decision.probabilities[2] });
  }
};

Open Weights Availability

Both models are available immediately on Hugging Face under the Apache 2.0 license:

  • cloudflare/clef-base (FP16 & INT8 quantized weights)
  • cloudflare/clef-flash (ONNX runtime optimized for edge deployment)

By providing open, sub-5ms probabilistic decision models, Cloudflare enables engineers to decouple fast routing and classification from expensive foundation model reasoning loops.