SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
61 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with SimpleBench
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About SimpleBench
What does SimpleBench measure?
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking. 61 AI models have been tested on it. Scores range from 1.4 to 78.3 out of 100.
Which model leads on SimpleBench?
Claude Fable 5 from Anthropic leads SimpleBench with a score of 78.3. The median score across 61 tested models is 29.4.
Is SimpleBench saturated?
No · the top score is 78.3 out of 100 (78%). There is still meaningful room for improvement on SimpleBench.
Does SimpleBench predict performance on other benchmarks?
Yes · SimpleBench scores correlate 1.00 with LiveBench · Reasoning across 5 shared models. Models that do well on SimpleBench tend to do well on LiveBench · Reasoning.
How often is SimpleBench data refreshed?
BenchGecko pulls updates daily. New model scores on SimpleBench appear as soon as they are published by Epoch AI or the model provider.
- Category
- Reasoning
- Max score
- 100
- Models
- 61
- Updated
- 2026-06-09
Top on SimpleBench
Claude Fable 5 · 78.3Gemini 3.1 Pro Preview · 75.5Gemini 3.5 Flash · 72.0Gemini 3 Pro · 71.7GPT-5.4 Pro · 68.9More reasoning benchmarks
Same category · related evaluations