Proofbench
The Frontier
Best score over time · one chart, every benchmark
Full rankings
56 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Proofbench
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Proofbench
What does Proofbench measure?
Proofbench is a knowledge benchmark in the BenchGecko catalog. 56 AI models have been tested on it. Scores range from 2.0 to 100.0 out of 100.
Which model leads on Proofbench?
Claude Fable 5.1 from Anthropic leads Proofbench with a score of 100.0. The median score across 56 tested models is 31.0.
Is Proofbench saturated?
Yes · the top model on Proofbench has reached 100.0 out of 100, within 5% of the theoretical ceiling. This benchmark is approaching saturation and may be replaced by a harder successor.
Does Proofbench predict performance on other benchmarks?
Yes · Proofbench scores correlate 0.98 with Osworld 2 0 across 7 shared models. Models that do well on Proofbench tend to do well on Osworld 2 0.
How often is Proofbench data refreshed?
BenchGecko pulls updates daily. New model scores on Proofbench appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 56
- Updated
- 2026-09-28
Top on Proofbench
Claude Fable 5.1 · 100.0Claude Opus 5.5 · 100.0Claude Sonnet 5.5 · 100.0Claude Opus 5 · 99.0GPT-6 Astra · 99.0More knowledge benchmarks
Same category · related evaluations