FrontierMath: The Benchmark for Evaluating Advanced Mathematical Reasoning
Authors: Epoch AI Research Team, Tamay Besiroglu, David Rein, Alexander Bates, Julian Michael
$\text{Verification: } \mathcal{V}(S_{\text{model}}, \text{SymPy/Lean Proof}) \in \{0, 1\}, \quad \text{Human PhD Baseline} \approx 80\%, \quad \text{Pre-2026 AI} < 5\%$
FrontierMath is specifically designed to eliminate data contamination and shallow memorization:
1. **Unpublished Research Problems**: All questions were authored specifically for the benchmark by active research mathematicians and kept strictly confidential.
2. **Semi-Automated Verifiability**: Problems are designed with unambiguous, computationally verifiable numerical, algebraic, or formal Lean theorem answers.
3. **Multi-Hour Problem Complexity**: Solving each question typically requires a human mathematics graduate student or professor several hours of concentrated intellectual effort.