Laya vs. Jev: How System 1 Decision Models Broke Free from the Cloud and Run in Your Browser
After TypeSafe AI’s Jev introduced fast non-autoregressive decisions, ConvAI Innovations open-sourced Laya on ModernBERT-large, and researcher Vishal Mysore got it running entirely inside web browsers via ONNX Runtime Web. An architectural breakdown of typed decision models, local WebGPU execution, JevBench benchmarks, and the end of text-generation overhead for software routing.
In mid-September 2026, the machine learning industry witnessed an unexpected convergence. For two years, software teams had relied on general autoregressive Large Language Models (LLMs)—such as GPT-4o, Claude 3.5 Sonnet, and Llama 3.3—to execute routine programmatic decisions: routing customer tickets, classifying intent, validating form inputs, and orchestrating agent pipelines.
Using a 70-billion-parameter text generator to select one choice from a dropdown menu created severe operational penalties: hundreds of milliseconds of token latency, random formatting drift, high cloud compute bills, and persistent privacy vulnerabilities.
The arrival of System 1 decision models solved this mismatch. On September 15, 2026, TypeSafe AI launched Jev, a hosted API designed to return calibrated probabilities for typed decisions in single-digit milliseconds. Three days later, on September 18, ConvAI Innovations open-sourced Laya, a 421-million-parameter Apache-2.0 model built on ModernBERT-large.
Shortly after, independent engineer and researcher Vishal Mysore (writing under the handle visrow) achieved a milestone: he converted Laya into ONNX format, quantized its weights, and launched layaForWeb—running the full decision model entirely within client-side web browsers without backend servers, API keys, or cloud latency.
1. The Core Paradigm: Generative LLMs vs. System 1 Decision Models
To understand why Jev went viral and why Laya in the browser matters, consider how a software program makes decisions:
Traditional LLM Pipeline (Autoregressive Text Generation):
[User Text] ──► [Cloud LLM (70B Params)] ──► Generates 80 tokens of markdown text ──► Regex parse ──► Boolean Action
Latency: 400ms - 1,800ms | Cost: High | Privacy: Zero | Reliability: Subject to JSON syntax drift
System 1 Decision Model Pipeline (Jev & Laya):
[User Text + Typed Questions] ──► [Decision Model (421M Params)] ──► Direct Probability Vector [0.94, 0.04, 0.02]
Latency: 12ms (Cloud Jev) / 28ms (In-Browser Laya) | Cost: Zero marginal | Privacy: Local | Reliability: 100% Deterministic
Daniel Kahneman's cognitive framework divides human thought into two modes:
- System 1: Fast, instinctive, non-deliberative pattern recognition.
- System 2: Slow, deliberative, sequential logical reasoning.
Standard autoregressive language models operate like System 2 engines forced to do System 1 labor. Every classification token requires a full forward pass through the transformer stack.
Decision models abandon text generation entirely. They take two inputs:
- State: The context or document under inspection (an email, error trace, clinical note, or database record).
- Typed Question: A structured question, which can take three forms:
- Choice: Selecting from a finite list of categories (e.g.,
["billing", "technical", "sales"]). - Score: Scoring an attribute along a defined scale (e.g., 1 to 5).
- Boolean: A binary confirmation (e.g.,
"Does this message contain personal identifying information?").
- Choice: Selecting from a finite list of categories (e.g.,
The model outputs a vector of raw float probabilities summing to 1.0. There are no intermediate conversational tokens, no markdown code fences (```json`), and no parser errors.
2. Jev vs. Laya: The Architectural Comparison
While both models target the System 1 decision tier, their business models, underlying neural architectures, and deployment profiles differ significantly.
| Technical Specification | TypeSafe AI Jev | ConvAI Innovations Laya | Laya In-Browser (layaForWeb) |
|---|---|---|---|
| Launch Date | September 15, 2026 | September 18, 2026 | September 2026 |
| Lead Architect / Creator | Diogo Almeida (ex-OpenAI) | ConvAI Innovations Core Team | Vishal Mysore (visrow) |
| Model Architecture | Proprietary Non-Autoregressive | ModernBERT-large Encoder Backbone | Quantized ModernBERT-large ONNX |
| Total Parameters | Proprietary (~350M - 500M est.) | 421 Million (395M ModernBERT) | 421 Million (Int8 / Int4) |
| License | Proprietary Hosted Cloud API | Apache 2.0 (Permissive Open Weights) | Apache 2.0 Open Source |
| Deployment Target | Cloud API / Cloudflare Workers AI | Self-Hosted Linux / vLLM / PyTorch | Client-Side WebGPU / WASM |
| Max Context Window | 64,000 tokens | 8,192 tokens | 4,096 - 8,192 tokens |
| Max Discrete Choices | Up to 255 options | 10 to 32 options recommended | 10 to 32 options |
| Zero-Shot Generalization | High (Broad out-of-the-box calibration) | Moderate (Excels with fine-tuning) | Moderate (Zero-shot client side) |
| Inference Latency | 8ms - 15ms (Dedicated Cloud) | 14ms - 22ms (Server GPU) | 25ms - 65ms (Client WebGPU/WASM) |
| Weight Footprint | Cloud-hosted | ~1.6 GB (FP16 PyTorch) | 422 MB (Int8) / 278 MB (Int4) |
| Infrastructure Cost | Pay-per-query ($0.06 - $0.12/1M) | Cloud GPU operational expense | $0.00 (Runs on user device) |
| Data Privacy | Cloud transmission (Encrypted) | Sovereign datacenter boundary | 100% Local (Data never leaves browser) |
3. How Vishal Mysore Brought Laya into the Web Browser
The breakthrough that captured the developer community was Vishal Mysore's demonstration: executing the complete 421M-parameter Laya decision model client-side inside standard browser tabs without contacting an origin server.
Browser Execution Architecture (layaForWeb):
┌────────────────────────────────────────────────────────────────────────┐
│ User Web Browser (Client) │
│ │
│ ┌────────────────────┐ ┌────────────────────────────────┐ │
│ │ Input Text & Form │ ───────► │ WebAssembly / WebGPU Context │ │
│ └────────────────────┘ │ (ONNX Runtime Web Runtime) │ │
│ └──────────────┬─────────────────┘ │
│ │ │
│ ▼ │
│ ┌────────────────────────────────┐ │
│ │ Quantized Laya Model │ │
│ │ - Int8: 422MB │ │
│ │ - Int4: 278MB │ │
│ └──────────────┬─────────────────┘ │
│ │ │
│ ▼ │
│ ┌────────────────────┐ ┌────────────────────────────────┐ │
│ │ Client UI Action │ ◄─────── │ Direct Probability Emission │ │
│ │ (Route, Gating) │ │ [Choice A: 0.88, B: 0.12] │ │
│ └────────────────────┘ └────────────────────────────────┘ │
└────────────────────────────────────────────────────────────────────────┘
▲
│ NO SERVER COMMUNICATED · 0 BYTES TRANSMITTED OVER WIRE
A. The Challenge of Client-Side Foundation Models
Standard open-weight models (such as Llama 3 8B or Mistral 7B) require 4GB to 16GB of VRAM. While WebLLM and transformers.js can run small generative language models in the browser, doing so consumes gigabytes of memory, heats up client processors, and takes tens of seconds to download.
Because Laya has 421 million parameters and uses an encoder-only architecture rather than an autoregressive decoder, its compute graph is compact and forward-only.
B. ONNX Quantization Strategy
Mysore converted the PyTorch weights to Open Neural Network Exchange (ONNX) format and implemented post-training quantization:
- FP16 Baseline: 842 MB. High fidelity, but heavy for initial browser asset caches.
- Int8 Quantization: 422 MB. Minimal loss in probability calibration across classification benchmarks.
- Int4 Quantization: 278 MB. Enables download and initial compilation over broadband in under five seconds, running inside standard browser tabs with modest memory footprints.
C. The WebGPU / WebAssembly Pipeline
Using Microsoft’s ONNX Runtime Web, the browser binds directly to the client’s GPU via the modern WebGPU standard. If the client machine lacks dedicated GPU acceleration or WebGPU flags, ONNX Runtime Web automatically falls back to multithreaded WebAssembly (WASM) with SIMD instruction sets.
D. Practical Code Implementation
Executing a typed decision in the browser requires straightforward JavaScript:
import * as ort from 'onnxruntime-web';
// Initialize the local in-browser session
const session = await ort.InferenceSession.create('/models/laya-int8.onnx', {
executionProviders: ['webgpu', 'wasm']
});
// Run a typed decision entirely client-side
export async function evaluateCustomerTicket(ticketText: string) {
const options = ['billing', 'bug_report', 'feature_request', 'security_issue'];
// Format input tensors for state and choice branches
const inputFeeds = formatLayaInputs(ticketText, options);
// Single forward pass: ~35ms on consumer Apple M-series or Intel Iris GPU
const output = await session.run(inputFeeds);
const probabilities = extractSoftmax(output.logits.data);
return {
bestMatch: options[probabilities.indexOf(Math.max(...probabilities))],
distribution: Object.fromEntries(options.map((opt, i) => [opt, probabilities[i]]))
};
}
4. ModernBERT: The Backbone Behind Laya
Laya’s performance is tied directly to its architectural foundation: ModernBERT-large.
For years, encoder-based models remained stuck in the 2018 era of BERT and RoBERTa, which lacked modern transformer advancements like FlashAttention, Rotary Position Embeddings (RoPE), and extended context handling. In late 2024 and early 2025, the research community released ModernBERT, redesigning encoder architectures from scratch.
ConvAI Innovations capitalized on these improvements when building Laya:
- Unpadding and Native Sequence Packing: Older encoders spent computation on padding tokens in short sentences. ModernBERT removes padding entirely from attention matrix multiplication, speeding up batch inference by 2.5x to 4x.
- Rotary Position Embeddings (RoPE): Replaces static sinusoidal embeddings with RoPE, expanding context support from BERT's old 512-token limit up to 8,192 tokens. Laya can process an entire 15-page legal contract or customer transcript in a single pass.
- GeGLU Activations & FlashAttention-2: Upgrades dense feed-forward networks with gated linear units and memory-efficient attention kernels, allowing Laya to run efficiently on small consumer hardware.
- Pre-training on Code and Text: ModernBERT was pre-trained on 2 trillion tokens of mixed programming code and web text, giving Laya strong comprehension of JSON schemas, stack traces, and technical payloads out of the box.
5. Benchmark Performance: JevBench v1.3.0 Analysis
To evaluate how Laya compares to Jev in real-world scenarios, we examine data from JevBench v1.3.0 (tested September 2026), an open benchmark that measures decision calibration, option scaling, and latency across enterprise classification tasks.
A. Accuracy and Calibration Across Option Counts
| Benchmark Evaluation Metric | TypeSafe AI Jev (Cloud API) | ConvAI Laya (FP16 PyTorch) | Laya in Browser (Int8 ONNX) | GPT-4o-mini (Few-Shot) |
|---|---|---|---|---|
| Binary Intent Accuracy (Yes/No) | 99.1% | 98.4% | 97.9% | 96.2% |
| 5-Way Category Classification | 96.8% | 94.6% | 93.8% | 91.5% |
| 20-Way Complex Routing | 93.2% | 87.4% | 85.9% | 82.4% |
| 100-Option Scaled Selection | 88.5% | 71.2% (Degrades) | 68.0% | 58.6% |
| Brier Calibration Score (Lower=Better) | 0.038 | 0.062 | 0.071 | 0.145 |
| Mean Execution Latency | 12.4 ms | 18.2 ms | 36.5 ms | 412.0 ms |
| Compute Cost per 100K Inferences | $1.20 | ~$0.85 (GPU Server) | $0.00 | $60.00 |
B. Understanding the Trade-offs
- Jev’s Superiority in Wide Option Sets: When an application requires selecting one outcome among 50 to 255 candidate tools or catalog items, Jev’s non-autoregressive decoding mechanism handles label entropy cleanly. Laya’s encoder heads begin to experience probability diffusion when option counts exceed 32.
- Laya’s Strength in Stable, High-Volume Triage: For routine production workloads (such as routing support tickets, content moderation, spam detection, or sentiment polarity), Laya achieves near-parity with Jev at zero API cost.
- Calibration Quality: Decision models must know when they are uncertain. Jev exhibits tight Brier score calibration out of the box. Laya's raw entropy can be noisy, making it advantageous to evaluate decisions using branch probability (the share of probability assigned to the selected route relative to alternatives) rather than raw internal confidence.
6. The Architecture of In-Browser Workflows: layaForWorkflows
Running decision models locally inside the browser has opened up new patterns for web architecture. Vishal Mysore’s companion project, layaForWorkflows, arranges in-browser decision models into state-graph workflows.
┌────────────────────────────────────────────────────────────────────────┐
│ IN-BROWSER STATE MACHINE WORKFLOW │
│ │
│ [Document Input] ──► [Node 1: Medical / Financial PII Check?] │
│ │ │
│ ┌────────────┴────────────┐ │
│ ▼ ▼ │
│ [Yes: P > 0.85] [No: P < 0.15] │
│ │ │ │
│ ▼ ▼ │
│ [Node 2: Mask Entities] [Node 3: Department Triage] │
│ (Local In-Browser WASM) (Local In-Browser Laya) │
│ │ │ │
│ ▼ ▼ │
│ [Save Encrypted to IndexedDB] [Dispatch Action] │
└────────────────────────────────────────────────────────────────────────┘
In this setup:
- Every node in the workflow is a typed decision or deterministic action.
- Data never leaves the client's device, enabling compliance with strict regulations like HIPAA, GDPR, and financial secrecy rules.
- Offline functionality: A user on a field laptop with zero internet connectivity can parse and triage hundreds of documents without service interruption.
7. Strategic Guide: When to Choose Jev vs. Laya
Engineering teams designing multi-tier agent pipelines should evaluate their operational constraints:
Choose TypeSafe AI Jev If:
- Broad Zero-Shot Surface: You need accurate classification across dynamic, frequently changing options without collecting training data or tuning weights.
- Large Contexts (>8K tokens): Your workflow evaluates extensive legal depositions, multi-file code diffs, or book chapters where Jev’s 64k context window is required.
- High Option Cardinality: Your system routes requests across 50 to 255 tool endpoints simultaneously.
- Turnkey Cloud Integration: Your stack already uses Vercel AI SDK, LangChain, or Cloudflare Workers AI and prioritizes managed API maintenance over self-hosting.
Choose ConvAI Laya / layaForWeb If:
- Zero Cloud Data Transmission: Customer privacy mandates that sensitive user inputs must not be sent to third-party model providers.
- In-Browser / Edge Operation: You are developing browser extensions, client-side web apps, or electron desktop software that must run without network access.
- Domain Fine-Tuning: You have 5,000 labeled enterprise tickets and want to fine-tune the ModernBERT backbone weights for 99.8% domain accuracy.
- Compute Budget Constraints: You process millions of operational decisions per day and want to avoid ongoing per-token API charges.
8. Looking Forward: The Decomposition of Generalist Models
The rapid adoption of Jev and Laya marks a broader transition in AI systems engineering: the decomposition of monolithic LLMs into modular components.
For the past three years, the industry applied general autoregressive models to every layer of software development. As production systems scale, engineers are reserving high-parameter reasoning models (like OpenAI o3, Claude 3.7 Sonnet, and Gemini 4 Pro) for complex analysis, while delegating routing, parsing, and triage to lightweight System 1 models.
With Laya running directly on client machines inside web browsers, the cost of deterministic AI decisions has dropped to zero. Software applications no longer need to pay API tolls or wait hundreds of milliseconds for a cloud LLM to decide what to do next.