Exploitbench
The Frontier
Best score over time · one chart, every benchmark
Full rankings
7 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 26.5 | |
| 2 | 26.1 | |
| 3 | 23.6 | |
| 4 | 18.4 | |
| 5 | 18.1 | |
| 6 | 13.7 | |
| 7 | 13.3 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Exploitbench
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Exploitbench
What does Exploitbench measure?
Exploitbench is a knowledge benchmark in the BenchGecko catalog. 7 AI models have been tested on it. Scores range from 13.3 to 26.5 out of 100.
Which model leads on Exploitbench?
Claude Opus 4.7 from Anthropic leads Exploitbench with a score of 26.5. The median score across 7 tested models is 18.4.
Is Exploitbench saturated?
No · the top score is 26.5 out of 100 (27%). There is still meaningful room for improvement on Exploitbench.
Does Exploitbench predict performance on other benchmarks?
Yes · Exploitbench scores correlate 0.99 with Lmca across 5 shared models. Models that do well on Exploitbench tend to do well on Lmca.
How often is Exploitbench data refreshed?
BenchGecko pulls updates daily. New model scores on Exploitbench appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 7
- Updated
- 2026-04-20
Top on Exploitbench
Claude Opus 4.7 · 26.5Gemini 3.1 Pro Preview · 26.1Claude Sonnet 4.6 · 23.6Kimi K2.6 · 18.4GLM 5.1 · 18.1More knowledge benchmarks
Same category · related evaluations