Artificial Analysis · SciCode
The Frontier
Best score over time · one chart, every benchmark
Full rankings
51 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Artificial Analysis · SciCode
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Artificial Analysis · SciCode
What does Artificial Analysis · SciCode measure?
Artificial Analysis · SciCode is a knowledge benchmark in the BenchGecko catalog. 51 AI models have been tested on it. Scores range from 14.4 to 66.9 out of 100.
Which model leads on Artificial Analysis · SciCode?
Claude Opus 5.5 from Anthropic leads Artificial Analysis · SciCode with a score of 66.9. The median score across 51 tested models is 45.5.
Is Artificial Analysis · SciCode saturated?
No · the top score is 66.9 out of 100 (67%). There is still meaningful room for improvement on Artificial Analysis · SciCode.
Does Artificial Analysis · SciCode predict performance on other benchmarks?
Yes · Artificial Analysis · SciCode scores correlate 0.94 with HLE across 5 shared models. Models that do well on Artificial Analysis · SciCode tend to do well on HLE.
How often is Artificial Analysis · SciCode data refreshed?
BenchGecko pulls updates daily. New model scores on Artificial Analysis · SciCode appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 51
- Updated
- 2026-09-29
Top on Artificial Analysis · SciCode
Claude Opus 5.5 · 66.9Claude Fable 5.1 · 63.1Claude Sonnet 5.5 · 61.0MiMo-V2.6-Pro · 60.9Kimi K3 · 59.5More knowledge benchmarks
Same category · related evaluations