FrontierMath-Tier-4-v2-Private
The Frontier
Best score over time · one chart, every benchmark
Full rankings
22 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 87.8 | |
| 2 | 56.1 | |
| 3 | 49.0 | |
| 4 | 34.1 | |
| 5 | 31.7 | |
| 6 | 31.7 | |
| 7 | 26.8 | |
| 8 | 26.8 | |
| 9 | 26.8 | |
| 10 | 25.6 | |
| 11 | 21.9 | |
| 12 | 19.5 | |
| 13 | 17.1 | |
| 14 | 12.2 | |
| 15 | 12.2 | |
| 16 | 12.2 | |
| 17 | 9.8 | |
| 18 | 4.9 | |
| 19 | 4.9 | |
| 20 | 2.4 | |
| 21 | 2.4 | |
| 22 | 2.4 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with FrontierMath-Tier-4-v2-Private
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About FrontierMath-Tier-4-v2-Private
What does FrontierMath-Tier-4-v2-Private measure?
FrontierMath-Tier-4-v2-Private is a knowledge benchmark in the BenchGecko catalog. 22 AI models have been tested on it. Scores range from 2.4 to 87.8 out of 100.
Which model leads on FrontierMath-Tier-4-v2-Private?
Claude Fable 5 from Anthropic leads FrontierMath-Tier-4-v2-Private with a score of 87.8. The median score across 22 tested models is 20.7.
Is FrontierMath-Tier-4-v2-Private saturated?
No · the top score is 87.8 out of 100 (88%). There is still meaningful room for improvement on FrontierMath-Tier-4-v2-Private.
Does FrontierMath-Tier-4-v2-Private predict performance on other benchmarks?
Yes · FrontierMath-Tier-4-v2-Private scores correlate 0.94 with FrontierMath-Tier-4-2025-07-01-Private across 19 shared models. Models that do well on FrontierMath-Tier-4-v2-Private tend to do well on FrontierMath-Tier-4-2025-07-01-Private.
How often is FrontierMath-Tier-4-v2-Private data refreshed?
BenchGecko pulls updates daily. New model scores on FrontierMath-Tier-4-v2-Private appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 22
- Updated
- 2026-06-12
Top on FrontierMath-Tier-4-v2-Private
Claude Fable 5 · 87.8Claude Opus 4.8 · 56.1GPT-5.4 · 49.0Qwen3.7 Max · 34.1Claude Opus 4.7 · 31.7More knowledge benchmarks
Same category · related evaluations