FrontierMath-2025-02-28-Private
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
65 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with FrontierMath-2025-02-28-Private
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About FrontierMath-2025-02-28-Private
What does FrontierMath-2025-02-28-Private measure?
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning. 65 AI models have been tested on it. Scores range from 0.3 to 50.0 out of 100.
Which model leads on FrontierMath-2025-02-28-Private?
GPT-5.4 Pro from OpenAI leads FrontierMath-2025-02-28-Private with a score of 50.0. The median score across 65 tested models is 12.4.
Is FrontierMath-2025-02-28-Private saturated?
No · the top score is 50.0 out of 100 (50%). There is still meaningful room for improvement on FrontierMath-2025-02-28-Private.
Does FrontierMath-2025-02-28-Private predict performance on other benchmarks?
Yes · FrontierMath-2025-02-28-Private scores correlate 0.99 with FrontierMath-Tiers-1-3-v2-Private across 20 shared models. Models that do well on FrontierMath-2025-02-28-Private tend to do well on FrontierMath-Tiers-1-3-v2-Private.
How often is FrontierMath-2025-02-28-Private data refreshed?
BenchGecko pulls updates daily. New model scores on FrontierMath-2025-02-28-Private appear as soon as they are published by Epoch AI or the model provider.
- Category
- Math
- Max score
- 100
- Models
- 65
- Updated
- 2026-05-27
Top on FrontierMath-2025-02-28-Private
GPT-5.4 Pro · 50.0GPT-5.4 · 47.6Claude Opus 4.8 · 47.2Claude Opus 4.7 · 43.8Claude Opus 4.6 · 40.7More math benchmarks
Same category · related evaluations