Balrog
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
27 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 58.1 | |
| 2 | 57.0 | |
| 3 | 48.1 | |
| 4 | 43.6 | |
| 5 | 43.5 | |
| 6 | 43.3 | |
| 7 | 34.9 | |
| 8 | 33.5 | |
| 9 | 32.8 | |
| 10 | 32.6 | |
| 11 | 32.3 | |
| 12 | 32.3 | |
| 13 | 31.2 | |
| 14 | 29.5 | |
| 15 | 27.9 | |
| 16 | 27.3 | |
| 17 | 23.0 | |
| 18 | 21.0 | |
| 19 | 21.0 | |
| 20 | 19.3 | |
| 21 | 17.6 | |
| 22 | 17.4 | |
| 23 | 17.4 | |
| 24 | 16.2 | |
| 25 | 15.1 | |
| 26 | 14.6 | |
| 27 | 11.6 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Balrog
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Balrog
What does Balrog measure?
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning. 27 AI models have been tested on it. Scores range from 11.6 to 58.1 out of 100.
Which model leads on Balrog?
Gemini 3 Pro from Google DeepMind leads Balrog with a score of 58.1. The median score across 27 tested models is 29.5.
Is Balrog saturated?
No · the top score is 58.1 out of 100 (58%). There is still meaningful room for improvement on Balrog.
Does Balrog predict performance on other benchmarks?
Yes · Balrog scores correlate 0.94 with Artificial Analysis · Quality Index across 5 shared models. Models that do well on Balrog tend to do well on Artificial Analysis · Quality Index.
How often is Balrog data refreshed?
BenchGecko pulls updates daily. New model scores on Balrog appear as soon as they are published by Epoch AI or the model provider.
- Category
- Reasoning
- Max score
- 100
- Models
- 27
- Updated
- 2026-02-19
Top on Balrog
Gemini 3 Pro · 58.1Gemini 3.1 Pro Preview · 57.0Gemini 3 Flash Preview · 48.1Grok 4 · 43.6Claude Opus 4.5 · 43.5More reasoning benchmarks
Same category · related evaluations