Osworld 2 0
The Frontier
Best score over time · one chart, every benchmark
Full rankings
8 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 31.4 | |
| 2 | 27.3 | |
| 3 | 20.6 | |
| 4 | 18.2 | |
| 5 | 9.3 | |
| 6 | 4.6 | |
| 7 | 4.6 | |
| 8 | 2.8 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Osworld 2 0
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Osworld 2 0
What does Osworld 2 0 measure?
Osworld 2 0 is a knowledge benchmark in the BenchGecko catalog. 8 AI models have been tested on it. Scores range from 2.8 to 31.4 out of 100.
Which model leads on Osworld 2 0?
Claude Opus 5 from Anthropic leads Osworld 2 0 with a score of 31.4. The median score across 8 tested models is 13.8.
Is Osworld 2 0 saturated?
No · the top score is 31.4 out of 100 (31%). There is still meaningful room for improvement on Osworld 2 0.
Does Osworld 2 0 predict performance on other benchmarks?
Yes · Osworld 2 0 scores correlate 0.99 with SimpleBench across 5 shared models. Models that do well on Osworld 2 0 tend to do well on SimpleBench.
How often is Osworld 2 0 data refreshed?
BenchGecko pulls updates daily. New model scores on Osworld 2 0 appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 8
- Updated
- 2026-07-24
Top on Osworld 2 0
Claude Opus 5 · 31.4GPT-5.6 Sol · 27.3Claude Opus 4.8 · 20.6Claude Opus 4.7 · 18.2Claude Sonnet 4.6 · 9.3More knowledge benchmarks
Same category · related evaluations