APEX-Agents
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
35 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 49.6 | |
| 2 | 45.0 | |
| 3 | 42.5 | |
| 4 | 35.9 | |
| 5 | 34.3 | |
| 6 | 33.9 | |
| 7 | 33.5 | |
| 8 | 31.7 | |
| 9 | 31.7 | |
| 10 | 24.6 | |
| 11 | 24.0 | |
| 12 | 23.7 | |
| 13 | 18.4 | |
| 14 | 18.4 | |
| 15 | 18.3 | |
| 16 | 17.5 | |
| 17 | 17.2 | |
| 18 | 17.2 | |
| 19 | 16.9 | |
| 20 | 15.2 | |
| 21 | 14.4 | |
| 22 | 13.6 | |
| 23 | 9.3 | |
| 24 | 8.9 | |
| 25 | 7.0 | |
| 26 | 6.6 | |
| 27 | 6.2 | |
| 28 | 4.7 | |
| 29 | 4.0 | |
| 30 | 3.1 | |
| 31 | 3.0 | |
| 32 | 2.1 | |
| 33 | 1.8 | |
| 34 | 1.1 | |
| 35 | 1.1 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with APEX-Agents
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About APEX-Agents
What does APEX-Agents measure?
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments. 35 AI models have been tested on it. Scores range from 1.1 to 49.6 out of 100.
Which model leads on APEX-Agents?
Gemini 3.5 Flash from Google DeepMind leads APEX-Agents with a score of 49.6. The median score across 35 tested models is 17.2.
Is APEX-Agents saturated?
No · the top score is 49.6 out of 100 (50%). There is still meaningful room for improvement on APEX-Agents.
Does APEX-Agents predict performance on other benchmarks?
Yes · APEX-Agents scores correlate 0.94 with Cybench across 5 shared models. Models that do well on APEX-Agents tend to do well on Cybench.
How often is APEX-Agents data refreshed?
BenchGecko pulls updates daily. New model scores on APEX-Agents appear as soon as they are published by Epoch AI or the model provider.
- Category
- Agent
- Max score
- 100
- Models
- 35
- Updated
- 2026-06-09
Top on APEX-Agents
Gemini 3.5 Flash · 49.6Claude Fable 5 · 45.0Claude Opus 4.8 · 42.5GPT-5.4 · 35.9GPT-5.2 · 34.3More agent benchmarks
Same category · related evaluations