Benchmark · KnowledgeSettled

Artificial Analysis · Coding Index

Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads.

Updated 2026-06-16

The composite index reveals that coding capability has become the most contested dimension among frontier models. The top-5 spread is consistently tighter than on any single coding benchmark alone.

Scoring: Weighted composite of multiple coding benchmark scores. Higher is better. Scale varies by evaluation cycle.

Models tested
89
Top score
76.5
Claude Fable 5
Median
35.5
min 0.8
Top-5 spread
σ 2.8
Competitive

Best score over time · one chart, every benchmark

ARTIFICIAL ANALYSIS · CODING INDEX87 MODELS · FRONTIER RUNNING MAX015304560SCORE ↑Jan 25May 25Sep 25Feb 26Jun 26RELEASE DATE →benchgecko.ai/benchmark/aa-coding-index · frontier
Frontier on Artificial Analysis · Coding Index rose from 11.2 to 76.5 in 17 months · +65.3 points · latest leader Claude Fable 5 from Anthropic.
Pink dots = frontier records · 11 totalClick to open model page

89 models tested · sorted by score

#ModelScore
1Anthropic logoClaude Fable 576.5
2Anthropic logoClaude Opus 4.8 (Fast)74.3
3Anthropic logoClaude Opus 4.7 (Fast)73.6
4Google DeepMind logoGemini 3.5 Flash70.1
5Google DeepMind logoGemini 3.1 Pro Preview68.8
6z-ai logoGLM 5.268.8
7Alibaba Qwen logoQwen3.7 Max66.0
8Anthropic logoClaude Sonnet 4.663.0
9moonshotai logoKimi K2.7 Code60.8
10xiaomi logoMiMo-V2.5-Pro60.2
11DeepSeek logoDeepSeek V4 Pro59.4
12
U
Muse Spark
58.6
13minimax logoMiniMax M358.6
14OpenAI logoGPT-5.457.3
15DeepSeek logoDeepSeek V4 Flash56.2
16OpenAI logoGPT-5.4 Mini56.1
17OpenAI logoGPT-5.4 Nano56.1
18moonshotai logoKimi K2.656.0
19Alibaba Qwen logoQwen3.7 Plus55.9
20z-ai logoGLM 5.155.8
21Alibaba Qwen logoQwen3.6 Plus54.5
22Alibaba Qwen logoQwen3.6 27B53.7
23OpenAI logoGPT-5.3-Codex53.1
24minimax logoMiniMax M2.752.6
25NVIDIA logoNemotron 3 Ultra49.3
26Alibaba Qwen logoQwen3.5 397B A17B48.2
27Anthropic logoClaude Opus 4.6 (Fast)48.1
28Mistral AI logoMistral Medium 3.546.9
29Alibaba Qwen logoQwen3.5-122B-A10B45.7
30z-ai logoGLM 5 Turbo44.2
31Google DeepMind logoGemma 4 31B (free)43.4
32
I
Ring-2.6-1T
42.8
33Google DeepMind logoGemini 3 Flash Preview42.6
34xAI logoGrok 4.342.3
35Alibaba Qwen logoQwen3.6 35B A3B41.9
36xiaomi logoMiMo-V2-Pro41.4
37moonshotai logoKimi K2.539.5
38Google DeepMind logoGemini 3 Pro39.4
39Google DeepMind logoGemma 4 26B A4B (free)39.3
40OpenAI logoo338.4
41DeepSeek logoDeepSeek V3.2 Speciale37.9
42stepfun logoStep 3.7 Flash37.3
43DeepSeek logoDeepSeek V3.236.7
44z-ai logoGLM 5V Turbo36.2
45xiaomi logoMiMo-V2-Omni35.5
46Alibaba Qwen logoQwen3.5-27B34.9
47Google DeepMind logoGemini 3.1 Flash Lite34.7
48xiaomi logoMiMo-V2-Flash33.5
49Google DeepMind logoGemini 2.5 Pro31.9
50stepfun logoStep 3.5 Flash31.6
51xAI logoGrok 4.1 Fast30.9
52inception logoMercury 230.6
53Alibaba Qwen logoQwen3 Max Thinking30.5
54OpenAI logogpt-oss-120b (free)30.4
55Alibaba Qwen logoQwen3.5-35B-A3B30.3
56Google DeepMind logoGemini 3.1 Flash Lite Preview30.1
57Alibaba Qwen logoQwen3.5-9B28.7
58arcee-ai logoTrinity Large Thinking27.2
59Alibaba Qwen logoQwen3 Coder 480B A35B (free)24.6
60Mistral AI logoMistral Small 424.3
61DeepSeek logoR1 052824.0
62xAI logoGrok Code Fast 123.7
63Alibaba Qwen logoQwen3 Coder Next22.9
64Alibaba logoQwen3.5 4B22.6
65OpenAI logogpt-oss-20b (free)20.7
66Alibaba logoQwen3.5 2B19.6
67Alibaba Qwen logoQwen3 Next 80B A3B Instruct (free)19.5
68prime-intellect logoINTELLECT-319.1
69Mistral AI logoMistral Medium 3.118.3
70Google DeepMind logoGemini 2.5 Flash Lite18.1
71Meta logoLlama 4 Maverick16.3
72upstage logoSolar Pro 316.2
73Alibaba Qwen logoQwen3 Next 80B A3B Instruct15.3
74Alibaba logoQwen3.5 0.8B15.0
75baidu logoERNIE 4.5 300B A47B 14.5
76NVIDIA logoLlama 3.1 Nemotron Ultra 253B v113.1
77DeepSeek logoR1 Distill Llama 70B11.4
78Microsoft logoPhi 411.2
79ibm-granite logoGranite 4.1 8B10.3
80Cohere logoCommand A9.9
81rekaai logoReka Flash 38.9
82
N
Nanbeige4.1 3B
8.9
83NVIDIA logoNVIDIA Nemotron Nano 9B V28.3
84Meta logoLlama 4 Scout8.2
85ibm-granite logoGranite 4.0 Micro5.0
86liquid logoLFM2-24B-A2B3.6
87Microsoft logoPhi 4 Mini Instruct3.6
88liquid logoLFM2.5-1.2B-Thinking (free)1.4
89liquid logoLFM2.5-1.2B-Instruct (free)0.8
Details
Category
Knowledge
Creator
Artificial Analysis
Max score
60
Modality
Code
Scoring
Weighted composite of multiple coding benchmark scores. Higher is better. Scale varies by evaluation cycle.
Models
89
Updated
2026-06-16
Tests
Code generationDebuggingMulti-language codingSoftware engineering
Does not test
VisionLong contextSafetyScientific reasoning
Gecko's Take

Composite indices are useful for quick comparison, but always drill into the components. A model with a high AA Coding Index might excel at HumanEval but struggle on SWE-bench. Trust the components for production decisions.

Same category · related evaluations