Research

Leaked Gemini 4 Pro Benchmarks Show Lead Over OpenAI and Anthropic Across DeepSWE, Terminal-Bench, and OSWorld

A leaked benchmark evaluation matrix pits Google DeepMind’s unreleased Gemini 4 Pro against Claude Opus 5.5 and GPT-6 Astra. The document reports scores of 88.7 on DeepSWE v1.1, 95.3 on Terminal-Bench 2.1, and 86.8 on OSWorld 2.0, alongside a 2-million-token context window and aggressive pricing of $2.25 input and $11.25 output per million tokens.

By FreakVinci · 2026-09-28 · 16 min read

An unverified evaluation chart circulating among AI researchers and developer channels shows Google DeepMind’s upcoming Gemini 4 Pro outperforming frontier models from OpenAI and Anthropic. The document measures Gemini 4 Pro against Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Astra across real-world software engineering, headless shell operations, and operating system agent workflows.

The leaked scorecard shows Gemini 4 Pro claiming top positions on three agentic benchmarks:

  • 88.7% on DeepSWE v1.1
  • 95.3% on Terminal-Bench 2.1
  • 86.8% on OSWorld 2.0

Alongside benchmark numbers, the sheet outlines technical specifications: a 2,000,000-token (2M) context window and an API rate of $2.25 per million input tokens and $11.25 per million output tokens.

AI tester Sree shared the document with an explicit warning that Google DeepMind has not verified the material. The disclosure has divided AI practitioners between anticipation of a competitive Google surge and skepticism regarding pre-release evaluation protocols.


1. The Leaked Evaluation Matrix

The document compares the three frontier systems across three primary evaluation regimes, context capacities, and API rate cards:

System / Specification Gemini 4 Pro (Leaked) Claude Opus 5.5 (Official) GPT-6 Astra (Official) Margin vs Closest Rival
Developer Google DeepMind Anthropic OpenAI —
DeepSWE v1.1 (Pass@1) 88.7% 84.2% 82.5% +4.5% vs Claude Opus 5.5
Terminal-Bench 2.1 95.3% 91.0% 89.4% +4.3% vs Claude Opus 5.5
OSWorld 2.0 (Success Rate) 86.8% 79.4% 78.1% +7.4% vs Claude Opus 5.5
Context Window 2,000,000 tokens 1,000,000 tokens 1,100,000 tokens +900,000 tokens (+81.8%)
Pricing: Input / 1M Tokens $2.25 $15.00 $5.00 55.0% lower than GPT-6 Astra
Pricing: Output / 1M Tokens $11.25 $75.00 $20.00 43.8% lower than GPT-6 Astra
Reported Verification Status Unverified Leak Published Technical Report Published Technical Report Pending Confirmation

The figures indicate consistent margins across code modification, shell environments, and graphical user interface tasks.


2. DeepSWE v1.1: Multi-Repository Code Resolution

SWE-bench Verified served as the software engineering standard throughout 2024 and 2025. In 2026, researchers adopted DeepSWE v1.1 to address benchmark saturation.

DeepSWE v1.1 requires models to resolve real GitHub pull requests across repositories with more than 100,000 lines of code. It enforces three constraints that earlier benchmarks omitted:

  1. Multi-file semantic dependencies: Fixes require modifying across multiple packages, build files, and configuration schemas simultaneously.
  2. Hidden regression test suites: The evaluation harness executes unit, integration, and performance regression suites that the model cannot inspect during generation.
  3. Reproducible reproduction scripts: The model must write an isolated reproduction test that fails on the unpatched codebase and passes after the patch.

Claude Opus 5.5 established an 84.2% baseline on DeepSWE v1.1 during its August release. GPT-6 Astra recorded 82.5%.

The leaked scorecard gives Gemini 4 Pro 88.7%, a 4.5 percentage point advantage over Anthropic.

Engineers attribute this reported difference to Google’s internal deployment of AlphaCode 3 verification routines during reinforcement learning from human and compiler feedback (RLCF). When generating candidate patches, the model executes lightweight AST-level validation loops to verify syntax, type safety, and import integrity before outputting the final git diff.

# DeepSWE v1.1 execution harness validation trace (illustrative flow)
git clone --depth 50 https://github.com/eval-org/target-runtime-suite
python -m deepswe.evaluate \
  --model "gemini-4-pro-checkpoint-0921" \
  --dataset "deepswe-v1.1-verified" \
  --test-mode "strict-regression" \
  --max-context 2000000 \
  --output-dir "./results/gemini4pro"

3. Terminal-Bench 2.1: Headless Command-Line Mastery

Autonomous coding models struggle when moving from passive code editing to active command-line execution. Terminal-Bench 2.1 evaluates this capability inside isolated Docker sandboxes.

The suite evaluates 420 multi-stage operations, including:

  • Compiling C++26 modules with specific compiler flags.
  • Resolving broken dynamic linker dependencies and shared object paths.
  • Diagnosing Linux network namespaces and firewall rules.
  • Migrating live database schemas under disk space constraints.
Benchmark Evaluation Task Gemini 4 Pro (Leaked) Claude Opus 5.5 GPT-6 Astra
Dynamic Linker Debugging (ld.so) 96.4% 92.1% 90.8%
Network Namespace Configuration 94.8% 89.7% 88.2%
Cross-Compilation Toolchains 95.8% 91.5% 89.9%
Multi-Step Shell Recovery 94.2% 90.6% 88.6%
Overall Terminal-Bench 2.1 Score 95.3% 91.0% 89.4%

Gemini 4 Pro’s reported 95.3% score reflects error-recovery behavior. Rather than repeating a failing command with trivial flag modifications, internal test logs describe the model inspecting /var/log, executing strace, and reading intermediate stderr buffers before re-attempting execution.


4. OSWorld 2.0: Multimodal Desktop Agent Autonomy

The largest numerical gap in the leaked document appears in OSWorld 2.0, an evaluation suite that tests autonomous control over an entire desktop environment via screen captures, mouse clicks, and keystrokes.

OSWorld 2.0 tasks involve cross-application operations across Ubuntu, macOS, and Windows:

  • Extracting tabular data from a PDF inside LibreOffice Calc, calculating net revenues, and pasting the chart into an email draft in Thunderbird.
  • Opening a video editing project, trimming specific clips using keyboard shortcuts, and exporting to WebM.
  • Configuring IDE settings across multiple virtual desktops.

Gemini 4 Pro reportedly achieved 86.8%, compared to 79.4% for Claude Opus 5.5 and 78.1% for GPT-6 Astra.

The 7.4 percentage point gap points to architectural advantages in Google DeepMind’s native multimodal processing. While rival agent scaffolds capture discrete screenshots every 500 milliseconds and feed them through separate vision encoders, Gemini 4 Pro processes continuous video-frame streams natively within its primary transformer attention layers. This continuous temporal understanding reduces missed UI states and latency during complex multi-click workflows.


5. Context Window and Hardware Infrastructure

Context handling remains a key differentiator. The leaked sheet confirms Gemini 4 Pro supports a production context window of 2,000,000 tokens.

System Context Window Needle-in-a-Haystack (2M) Primary Compute Substrate
Gemini 4 Pro 2,000,000 tokens 99.8% Retrieval Accuracy TPU v6 (Trillium)
Claude Opus 5.5 1,000,000 tokens Not Supported (>1M) AWS Trainium 2 / NVIDIA H200
GPT-6 Astra 1,100,000 tokens Not Supported (>1.1M) NVIDIA B200 (Blackwell)

A 2M context window allows an engineering team to load an entire application repository, along with dependency documentation, commit history, and system logs, into a single prompt without chunking or vector search approximation.

Google achieves this throughput via its TPU v6 Trillium infrastructure. Trillium pods use optical circuit switches (OCS) to link thousands of tensor processing units with low interconnect latency, running paged KV-cache compression algorithms that cut memory overhead during multi-million token inference.


6. The Rate Card: Price Disruption in Frontier APIs

The leaked pricing details indicate an aggressive strategy by Google to drive developer adoption on Vertex AI and Google AI Studio:

Model Tier Input / Million Tokens Output / Million Tokens Blended Cost (3:1 Ratio)
Gemini 4 Pro (Leaked) $2.25 $11.25 $4.50
GPT-6 Astra $5.00 $20.00 $8.75
Claude Opus 5.5 $15.00 $75.00 $30.00
GPT-6 Sol $2.50 $10.00 $4.38
Claude Sonnet 5 $3.00 $15.00 $6.00

At $2.25 input and $11.25 output per million tokens, Gemini 4 Pro costs less than one-sixth of Claude Opus 5.5, while beating it across all three evaluation suites in the leaked test.

If these rates materialize at launch, Google will position Gemini 4 Pro against mid-tier reasoning models like GPT-6 Sol and Claude Sonnet 5 on price, while competing with top-tier flagship models on performance.


7. Leak Origin, Sree's Disclosures, and Community Skepticism

The benchmark sheet circulated publicly after AI tester Sree shared it across developer networks. Sree accompanied the graphic with clear cautionary guidance, emphasizing that no official party has confirmed the scores.

The AI research community reacted with a mixture of interest and skepticism:

Arguments Supporting the Leak

  • Evaluation Consistency: The 88.7% DeepSWE and 86.8% OSWorld numbers align with performance curves extrapolated from Gemini 2.5 Flash and DeepMind's AlphaProof research papers published earlier in 2026.
  • Infrastructure Timeline: Google began deploying TPU v6 Trillium clusters in late 2025. A late-September training checkpoint ready for October deployment matches Google’s hardware cycle.
  • Pricing Pattern: Google used aggressive pricing during the Gemini 1.5 rollout to capture market share from OpenAI. A $2.25/$11.25 structure matches Google’s cloud customer acquisition goals.

Arguments for Skepticism

  • Evaluation Scaffolding Discrepancies: The leaked sheet does not define the execution parameters used for the evaluation. An 88.7% score achieved via high-cost multi-path tree-of-thought sampling (pass@8) differs from a standard single-pass (pass@1) test.
  • Potential Prompt Contamination: Pre-release models evaluated on public benchmarks can exhibit synthetic overlap if the training data scraper ingested evaluation repositories or solution discussions.
  • Alpha Checkpoint Cherry-Picking: Internal research teams frequently run dozens of experimental post-training runs. Scores from an experimental branch may not survive the safety, refusal, and alignment passes required for commercial release.

8. What Happens Next: The October Launch Window

Google has not issued a formal statement regarding the leak. Historically, Google DeepMind ignores unverified leaks until official launch keynotes.

Industry watchers expect Google to unveil Gemini 4 Pro during an early October broadcast, integrating the model into Google AI Studio, Vertex AI, and Gemini Advanced. If the final production weights match the numbers in Sree’s leaked chart, the competitive pressure will shift to OpenAI and Anthropic to accelerate their late-2026 releases.