Artificial Analysis · Long Context Reasoning
The Frontier
Best score over time · one chart, every benchmark
Full rankings
61 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Artificial Analysis · Long Context Reasoning
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Artificial Analysis · Long Context Reasoning
What does Artificial Analysis · Long Context Reasoning measure?
Artificial Analysis · Long Context Reasoning is a knowledge benchmark in the BenchGecko catalog. 61 AI models have been tested on it. Scores range from 5.7 to 88.7 out of 100.
Which model leads on Artificial Analysis · Long Context Reasoning?
Kimi K3 from moonshotai leads Artificial Analysis · Long Context Reasoning with a score of 88.7. The median score across 61 tested models is 71.7.
Is Artificial Analysis · Long Context Reasoning saturated?
No · the top score is 88.7 out of 100 (89%). There is still meaningful room for improvement on Artificial Analysis · Long Context Reasoning.
Does Artificial Analysis · Long Context Reasoning predict performance on other benchmarks?
Yes · Artificial Analysis · Long Context Reasoning scores correlate 0.92 with Artificial Analysis · MMMU Pro across 31 shared models. Models that do well on Artificial Analysis · Long Context Reasoning tend to do well on Artificial Analysis · MMMU Pro.
How often is Artificial Analysis · Long Context Reasoning data refreshed?
BenchGecko pulls updates daily. New model scores on Artificial Analysis · Long Context Reasoning appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 61
- Updated
- 2026-09-29
Top on Artificial Analysis · Long Context Reasoning
Kimi K3 · 88.7MiMo-V2.6-Pro · 86.3Claude Fable 5.1 · 85.3Claude Opus 5.5 · 84.7DeepSeek V4.1 Flash · 84.0More knowledge benchmarks
Same category · related evaluations