Artificial Analysis · Terminal-Bench Hard
The Frontier
Best score over time · one chart, every benchmark
Full rankings
32 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 57.6 | |
| 2 | 53.8 | |
| 3 | 53.0 | |
| 4 | 47.0 | |
| 5 | 44.7 | |
| 6 | 42.4 | |
| 7 | 37.1 | |
| 8 | 36.4 | |
| 9 | 36.4 | |
| 10 | 35.6 | |
| 11 | 34.8 | |
| 12 | 33.3 | |
| 13 | 31.1 | |
| 14 | 29.5 | |
| 15 | 28.8 | |
| 16 | 25.0 | |
| 17 | 23.5 | |
| 18 | 22.7 | |
| 19 | 18.2 | |
| 20 | 18.2 | |
| 21 | 18.2 | |
| 22 | 17.4 | |
| 23 | 13.6 | |
| 24 | 10.6 | |
| 25 | 9.8 | |
| 26 | 6.8 | |
| 27 | 3.8 | |
| 28 | 3.8 | |
| 29 | 1.5 | |
| 30 | 1.5 | |
| 31 | 1.5 | |
| 32 | 0.8 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Artificial Analysis · Terminal-Bench Hard
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Artificial Analysis · Terminal-Bench Hard
What does Artificial Analysis · Terminal-Bench Hard measure?
Artificial Analysis · Terminal-Bench Hard is a knowledge benchmark in the BenchGecko catalog. 32 AI models have been tested on it. Scores range from 0.8 to 57.6 out of 100.
Which model leads on Artificial Analysis · Terminal-Bench Hard?
GPT-5.6 Terra from OpenAI leads Artificial Analysis · Terminal-Bench Hard with a score of 57.6. The median score across 32 tested models is 24.3.
Is Artificial Analysis · Terminal-Bench Hard saturated?
No · the top score is 57.6 out of 100 (58%). There is still meaningful room for improvement on Artificial Analysis · Terminal-Bench Hard.
Does Artificial Analysis · Terminal-Bench Hard predict performance on other benchmarks?
Yes · Artificial Analysis · Terminal-Bench Hard scores correlate 1.00 with ARC-AGI across 5 shared models. Models that do well on Artificial Analysis · Terminal-Bench Hard tend to do well on ARC-AGI.
How often is Artificial Analysis · Terminal-Bench Hard data refreshed?
BenchGecko pulls updates daily. New model scores on Artificial Analysis · Terminal-Bench Hard appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 32
- Updated
- 2026-09-22
Top on Artificial Analysis · Terminal-Bench Hard
GPT-5.6 Terra · 57.6Gemini 3.1 Pro Preview · 53.8GPT-5.3-Codex · 53.0Qwen3.7 Plus · 47.0Kimi K2.7 Code · 44.7More knowledge benchmarks
Same category · related evaluations