FrontierMath-Tiers-1-3-v2-Private
The Frontier
Best score over time · one chart, every benchmark
Full rankings
24 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 87.0 | |
| 2 | 80.0 | |
| 3 | 78.6 | |
| 4 | 70.2 | |
| 5 | 67.4 | |
| 6 | 66.0 | |
| 7 | 64.6 | |
| 8 | 62.8 | |
| 9 | 59.6 | |
| 10 | 57.2 | |
| 11 | 55.8 | |
| 12 | 55.4 | |
| 13 | 54.0 | |
| 14 | 51.2 | |
| 15 | 51.2 | |
| 16 | 46.7 | |
| 17 | 44.9 | |
| 18 | 36.1 | |
| 19 | 34.4 | |
| 20 | 24.6 | |
| 21 | 23.9 | |
| 22 | 20.0 | |
| 23 | 18.6 | |
| 24 | 12.6 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with FrontierMath-Tiers-1-3-v2-Private
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About FrontierMath-Tiers-1-3-v2-Private
What does FrontierMath-Tiers-1-3-v2-Private measure?
FrontierMath-Tiers-1-3-v2-Private is a knowledge benchmark in the BenchGecko catalog. 24 AI models have been tested on it. Scores range from 12.6 to 87.0 out of 100.
Which model leads on FrontierMath-Tiers-1-3-v2-Private?
Claude Fable 5 from Anthropic leads FrontierMath-Tiers-1-3-v2-Private with a score of 87.0. The median score across 24 tested models is 54.7.
Is FrontierMath-Tiers-1-3-v2-Private saturated?
No · the top score is 87.0 out of 100 (87%). There is still meaningful room for improvement on FrontierMath-Tiers-1-3-v2-Private.
Does FrontierMath-Tiers-1-3-v2-Private predict performance on other benchmarks?
Yes · FrontierMath-Tiers-1-3-v2-Private scores correlate 0.99 with FrontierMath-2025-02-28-Private across 20 shared models. Models that do well on FrontierMath-Tiers-1-3-v2-Private tend to do well on FrontierMath-2025-02-28-Private.
How often is FrontierMath-Tiers-1-3-v2-Private data refreshed?
BenchGecko pulls updates daily. New model scores on FrontierMath-Tiers-1-3-v2-Private appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 24
- Updated
- 2026-06-12
Top on FrontierMath-Tiers-1-3-v2-Private
Claude Fable 5 · 87.0Claude Opus 4.8 · 80.0GPT-5.4 · 78.6Claude Opus 4.7 · 70.2GPT-5.2 · 67.4More knowledge benchmarks
Same category · related evaluations