FrontierMath-Tier-4-2025-07-01-Private
FrontierMath Tier 4 (Jul 2025) · the most challenging tier of frontier mathematics, containing problems that push the absolute limits of AI mathematical reasoning.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
40 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 37.5 | |
| 2 | 31.3 | |
| 3 | 31.3 | |
| 4 | 27.1 | |
| 5 | 22.9 | |
| 6 | 22.9 | |
| 7 | 18.8 | |
| 8 | 18.8 | |
| 9 | 16.7 | |
| 10 | 14.6 | |
| 11 | U Muse Spark | 14.6 |
| 12 | 14.6 | |
| 13 | 14.6 | |
| 14 | 12.5 | |
| 15 | 12.5 | |
| 16 | 12.5 | |
| 17 | 8.3 | |
| 18 | 8.3 | |
| 19 | 6.3 | |
| 20 | 6.3 | |
| 21 | 6.3 | |
| 22 | 4.2 | |
| 23 | 4.2 | |
| 24 | 4.2 | |
| 25 | 4.2 | |
| 26 | 4.2 | |
| 27 | 4.2 | |
| 28 | 4.2 | |
| 29 | 4.2 | |
| 30 | 4.2 | |
| 31 | 4.2 | |
| 32 | 2.1 | |
| 33 | 2.1 | |
| 34 | 2.1 | |
| 35 | 2.1 | |
| 36 | 2.1 | |
| 37 | 2.1 | |
| 38 | 2.1 | |
| 39 | 2.1 | |
| 40 | 2.1 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with FrontierMath-Tier-4-2025-07-01-Private
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About FrontierMath-Tier-4-2025-07-01-Private
What does FrontierMath-Tier-4-2025-07-01-Private measure?
FrontierMath Tier 4 (Jul 2025) · the most challenging tier of frontier mathematics, containing problems that push the absolute limits of AI mathematical reasoning. 40 AI models have been tested on it. Scores range from 2.1 to 37.5 out of 100.
Which model leads on FrontierMath-Tier-4-2025-07-01-Private?
GPT-5.4 Pro from OpenAI leads FrontierMath-Tier-4-2025-07-01-Private with a score of 37.5. The median score across 40 tested models is 6.3.
Is FrontierMath-Tier-4-2025-07-01-Private saturated?
No · the top score is 37.5 out of 100 (38%). There is still meaningful room for improvement on FrontierMath-Tier-4-2025-07-01-Private.
Does FrontierMath-Tier-4-2025-07-01-Private predict performance on other benchmarks?
Yes · FrontierMath-Tier-4-2025-07-01-Private scores correlate 0.94 with FrontierMath-Tier-4-v2-Private across 19 shared models. Models that do well on FrontierMath-Tier-4-2025-07-01-Private tend to do well on FrontierMath-Tier-4-v2-Private.
How often is FrontierMath-Tier-4-2025-07-01-Private data refreshed?
BenchGecko pulls updates daily. New model scores on FrontierMath-Tier-4-2025-07-01-Private appear as soon as they are published by Epoch AI or the model provider.
- Category
- Math
- Max score
- 100
- Models
- 40
- Updated
- 2026-05-27
Top on FrontierMath-Tier-4-2025-07-01-Private
GPT-5.4 Pro · 37.5GPT-5.2 Pro · 31.3Claude Opus 4.8 · 31.3GPT-5.4 · 27.1Claude Opus 4.7 · 22.9More math benchmarks
Same category · related evaluations