Frontier AI Model Benchmark Matrix

Interactive, peer-reviewed evaluation of frontier AI architectures and reasoning models.

Verified Frontier AI Model Leaderboard

Empirical benchmark evaluation across coding autonomy (SWE-bench), reasoning (MMLU-Pro), and token economics.

Model Lab SWE-bench MMLU-Pro Throughput 1M Tokens
Gemini 4 Argon Google DeepMind 78.4% 93.6% 125 tps $1.20
GPT-6.1 Sol OpenAI 74.8% 92.4% 120 tps $1.25
Gemini 4 Pro (Leaked Benchmark) Google DeepMind 88.7% 93.8% 115 tps $2.25
Space Bunny Alpha (Stealth) MiniMax (Unannounced Stealth Release via OpenRouter) 65.4% 86.2% 110 tps Free
GPT-6 Sol OpenAI 68.8% 90.8% 94 tps $2.00
GPT-6 Luna OpenAI 48.2% 74.6% 157 tps $0.10
Claude Opus 5.5 Anthropic 89.9% 92.4% 82 tps $4.00
Claude Sonnet 5.5 Anthropic 82.4% 88.4% 112 tps $2.00
Manus 2.0 Manus AI 74.8% 85.6% 90 tps $3.50
ElevenLabs Eleven V4 & V4 Turbo ElevenLabs — — 240 tps $0.15
Qwen 4 Max (Leaked Staging) Alibaba Cloud 78.6% 91.4% 96 tps $1.20
DeepSeek V5 (Leaked Staging) DeepSeek AI 81.4% 87.5% 125 tps $0.12
MiniMax M3.1-Flash-Preview MiniMax 73.8% 81.2% 165 tps $0.10
Union Alpha (Stealth) OpenRouter & OpenCode 74.2% 89.8% 68 tps Free
Muse Spark 1.3 (xhigh) Meta AI 75.4% 88.4% 92 tps $1.25