Artificial Analysis · Coding Index
Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads.
The composite index reveals that coding capability has become the most contested dimension among frontier models. The top-5 spread is consistently tighter than on any single coding benchmark alone.
Scoring: Weighted composite of multiple coding benchmark scores. Higher is better. Scale varies by evaluation cycle.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
89 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Artificial Analysis · Coding Index
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
How it works
Evaluation methodology
The AA Coding Index is a composite score computed by Artificial Analysis that aggregates results from multiple coding benchmarks into a single normalized index. The component benchmarks include HumanEval, SWE-bench variants, code generation, debugging, and multi-language coding tasks. Each component is weighted and normalized to produce a comparable score across models. The exact weighting formula is proprietary to Artificial Analysis but designed to reflect real-world coding utility.
Industry relevance
Why teams track this benchmark
Individual coding benchmarks each test one facet of programming ability. The AA Coding Index compresses them into a single number that answers: "How good is this model at coding, overall?" For engineering teams evaluating models, this saves running five separate benchmarks.
Practical takeaways
By role
Use the AA Coding Index for initial model shortlisting, then run your own SWE-bench or Aider evaluation on your specific codebase for final selection.
The coding index correlates strongly with developer adoption. Models that lead here tend to capture the developer tooling market within 2-3 months.
The composite nature makes this less useful for ablation studies. Use individual component benchmarks for controlled experiments.
Frequently asked
About Artificial Analysis · Coding Index
What does Artificial Analysis · Coding Index measure?
Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads. 89 AI models have been tested on it. Scores range from 0.8 to 76.5 out of 60.
Which model leads on Artificial Analysis · Coding Index?
Claude Fable 5 from Anthropic leads Artificial Analysis · Coding Index with a score of 76.5. The median score across 89 tested models is 35.5.
Is Artificial Analysis · Coding Index saturated?
Yes · the top model on Artificial Analysis · Coding Index has reached 76.5 out of 60, within 5% of the theoretical ceiling. This benchmark is approaching saturation and may be replaced by a harder successor.
Does Artificial Analysis · Coding Index predict performance on other benchmarks?
Yes · Artificial Analysis · Coding Index scores correlate 0.96 with HELM · WildBench across 5 shared models. Models that do well on Artificial Analysis · Coding Index tend to do well on HELM · WildBench.
How often is Artificial Analysis · Coding Index data refreshed?
BenchGecko pulls updates daily. New model scores on Artificial Analysis · Coding Index appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Creator
- Artificial Analysis
- Max score
- 60
- Modality
- Code
- Scoring
- Weighted composite of multiple coding benchmark scores. Higher is better. Scale varies by evaluation cycle.
- Models
- 89
- Updated
- 2026-06-16
“Composite indices are useful for quick comparison, but always drill into the components. A model with a high AA Coding Index might excel at HumanEval but struggle on SWE-bench. Trust the components for production decisions.”
Top on Artificial Analysis · Coding Index
Claude Fable 5 · 76.5Claude Opus 4.8 (Fast) · 74.3Claude Opus 4.7 (Fast) · 73.6Gemini 3.5 Flash · 70.1Gemini 3.1 Pro Preview · 68.8More knowledge benchmarks
Same category · related evaluations