Artificial Analysis · Humanity's Last Exam
The Frontier
Best score over time · one chart, every benchmark
Full rankings
66 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Artificial Analysis · Humanity's Last Exam
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Artificial Analysis · Humanity's Last Exam
What does Artificial Analysis · Humanity's Last Exam measure?
Artificial Analysis · Humanity's Last Exam is a knowledge benchmark in the BenchGecko catalog. 66 AI models have been tested on it. Scores range from 1.1 to 61.4 out of 100.
Which model leads on Artificial Analysis · Humanity's Last Exam?
Claude Opus 5.5 from Anthropic leads Artificial Analysis · Humanity's Last Exam with a score of 61.4. The median score across 66 tested models is 20.0.
Is Artificial Analysis · Humanity's Last Exam saturated?
No · the top score is 61.4 out of 100 (61%). There is still meaningful room for improvement on Artificial Analysis · Humanity's Last Exam.
Does Artificial Analysis · Humanity's Last Exam predict performance on other benchmarks?
Yes · Artificial Analysis · Humanity's Last Exam scores correlate 0.99 with Terminal Bench across 6 shared models. Models that do well on Artificial Analysis · Humanity's Last Exam tend to do well on Terminal Bench.
How often is Artificial Analysis · Humanity's Last Exam data refreshed?
BenchGecko pulls updates daily. New model scores on Artificial Analysis · Humanity's Last Exam appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 66
- Updated
- 2026-09-29
Top on Artificial Analysis · Humanity's Last Exam
Claude Opus 5.5 · 61.4Claude Fable 5.1 · 59.1Claude Sonnet 5.5 · 55.0GPT-6 Astra · 54.7GPT-6.1 Sol · 52.9More knowledge benchmarks
Same category · related evaluations