Claude Fable 5.5 (Leaked Checkpoint) vs GPT-Image-2.5 (Sunburst)

A direct, empirical evaluation of Claude Fable 5.5 (Leaked Checkpoint) (Anthropic) against GPT-Image-2.5 (Sunburst) (OpenAI).

Claude Fable 5.5 (Leaked Checkpoint)

Developed by Anthropic · Proprietary Frontier Dynamic Reasoner parameters

Terminal-Bench 4.0 autonomous shell execution (78.5% target), recursive self-critique, and 1.5M-token repository refactoring.

GPT-Image-2.5 (Sunburst)

Developed by OpenAI · Proprietary Multimodal DiT parameters

High-fidelity typographic rendering, raytracing reflection accuracy, sketch-to-image conditioning, and localized comment inpainting.

Metric Claude Fable 5.5 (Leaked Checkpoint) GPT-Image-2.5 (Sunburst)
SWE-bench Verified 84.2% 62.4%
MMLU-Pro 94.6% 89.5%
MATH-500 98.4% 91.2%
Context Window 1,500,000 tokens Spatial Vector Token Canvas (1792×1024)
Input Pricing (per 1M) $3.50 $2.50
Output Pricing (per 1M) $17.50 $10.00