Metr Time Horizons
The Frontier
Best score over time · one chart, every benchmark
Full rankings
29 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 78.9 | |
| 2 | 77.0 | |
| 3 | 75.3 | |
| 4 | 75.0 | |
| 5 | 74.5 | |
| 6 | 74.3 | |
| 7 | 71.0 | |
| 8 | 69.6 | |
| 9 | 67.4 | |
| 10 | 66.8 | |
| 11 | 66.6 | |
| 12 | 65.4 | |
| 13 | 63.9 | |
| 14 | 63.9 | |
| 15 | 62.0 | |
| 16 | 60.0 | |
| 17 | 59.2 | |
| 18 | 56.6 | |
| 19 | 55.4 | |
| 20 | 51.9 | |
| 21 | 51.0 | |
| 22 | 47.4 | |
| 23 | 45.1 | |
| 24 | 40.1 | |
| 25 | 35.8 | |
| 26 | 33.8 | |
| 27 | 29.9 | |
| 28 | 29.5 | |
| 29 | 29.3 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Metr Time Horizons
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Metr Time Horizons
What does Metr Time Horizons measure?
Metr Time Horizons is a knowledge benchmark in the BenchGecko catalog. 29 AI models have been tested on it. Scores range from 29.3 to 78.9 out of 100.
Which model leads on Metr Time Horizons?
Claude Opus 4.6 from Anthropic leads Metr Time Horizons with a score of 78.9. The median score across 29 tested models is 62.0.
Is Metr Time Horizons saturated?
No · the top score is 78.9 out of 100 (79%). There is still meaningful room for improvement on Metr Time Horizons.
Does Metr Time Horizons predict performance on other benchmarks?
Yes · Metr Time Horizons scores correlate 0.97 with SWE-Bench verified across 15 shared models. Models that do well on Metr Time Horizons tend to do well on SWE-Bench verified.
How often is Metr Time Horizons data refreshed?
BenchGecko pulls updates daily. New model scores on Metr Time Horizons appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 29
- Updated
- 2026-03-05
Top on Metr Time Horizons
Claude Opus 4.6 · 78.9Gemini 3.1 Pro Preview · 77.0GPT-5.2 · 75.3Claude Opus 4.5 · 75.0GPT-5.3-Codex · 74.5More knowledge benchmarks
Same category · related evaluations