Deepswe
The Frontier
Best score over time · one chart, every benchmark
Full rankings
23 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 74.1 | |
| 2 | 73.8 | |
| 3 | 73.7 | |
| 4 | 72.7 | |
| 5 | 69.9 | |
| 6 | 69.6 | |
| 7 | 69.0 | |
| 8 | 68.5 | |
| 9 | 67.5 | |
| 10 | 67.2 | |
| 11 | 65.5 | |
| 12 | 63.4 | |
| 13 | 59.0 | |
| 14 | 57.5 | |
| 15 | 53.9 | |
| 16 | 53.8 | |
| 17 | 51.8 | |
| 18 | 46.7 | |
| 19 | 43.8 | |
| 20 | 37.4 | |
| 21 | 30.5 | |
| 22 | 29.9 | |
| 23 | 11.7 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Deepswe
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Deepswe
What does Deepswe measure?
Deepswe is a knowledge benchmark in the BenchGecko catalog. 23 AI models have been tested on it. Scores range from 11.7 to 74.1 out of 100.
Which model leads on Deepswe?
GPT-6 Astra from OpenAI leads Deepswe with a score of 74.1. The median score across 23 tested models is 63.4.
Is Deepswe saturated?
No · the top score is 74.1 out of 100 (74%). There is still meaningful room for improvement on Deepswe.
Does Deepswe predict performance on other benchmarks?
Yes · Deepswe scores correlate 0.88 with Artificial Analysis · GDPval across 9 shared models. Models that do well on Deepswe tend to do well on Artificial Analysis · GDPval.
How often is Deepswe data refreshed?
BenchGecko pulls updates daily. New model scores on Deepswe appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 23
- Updated
- 2026-09-04
Top on Deepswe
GPT-6 Astra · 74.1Gemini 3.8 Flash · 73.8Claude Opus 5 · 73.7GPT-5.6 Sol · 72.7Claude Fable 5 · 69.9More knowledge benchmarks
Same category · related evaluations