Benchmark · KnowledgeSettled

Artificial Analysis · Agentic Index

Artificial Analysis Agentic Index · a composite score measuring how well a model performs in agentic workflows · multi-step tool use, planning, error recovery, and autonomous task completion. Aggregates results from multiple agentic benchmarks including SWE-bench, tool-use tests, and planning evaluations. The canonical single-number metric for "how good is this model as an agent?"

Updated 2026-06-16

The agentic index shows the fastest frontier progression of any composite metric. Models released 6 months apart show 15-20 point gaps, reflecting rapid improvements in tool use and planning.

Scoring: Weighted composite of multiple agentic benchmark scores. Components include SWE-bench, tool use, planning, and error recovery. Higher is better.

Models tested
84
Top score
69.4
GPT-5.4
Median
27.5
min 1.1
Top-5 spread
σ 2.9
Competitive

Best score over time · one chart, every benchmark

ARTIFICIAL ANALYSIS · AGENTIC INDEX82 MODELS · FRONTIER RUNNING MAX015304560SCORE ↑Feb 25Jun 25Oct 25Feb 26Jun 26RELEASE DATE →benchgecko.ai/benchmark/aa-agentic-index · frontier
Frontier on Artificial Analysis · Agentic Index rose from 2.7 to 69.4 in 13 months · +66.7 points · latest leader GPT-5.4 from OpenAI.
Pink dots = frontier records · 8 totalClick to open model page

84 models tested · sorted by score

#ModelScore
1OpenAI logoGPT-5.469.4
2Anthropic logoClaude Opus 4.6 (Fast)67.6
3z-ai logoGLM 5 Turbo63.1
4xiaomi logoMiMo-V2-Pro62.8
5OpenAI logoGPT-5.3-Codex62.2
6z-ai logoGLM 5V Turbo61.1
7moonshotai logoKimi K2.558.9
8xiaomi logoMiMo-V2-Omni58.6
9Alibaba Qwen logoQwen3.5-27B54.6
10DeepSeek logoDeepSeek V3.252.9
11Anthropic logoClaude Fable 552.8
12stepfun logoStep 3.5 Flash52.0
13Alibaba Qwen logoQwen3 Max Thinking50.1
14Google DeepMind logoGemini 3 Flash Preview49.7
15xAI logoGrok 4.1 Fast49.3
16xiaomi logoMiMo-V2-Flash48.8
17Anthropic logoClaude Opus 4.8 (Fast)47.2
18Google DeepMind logoGemini 3 Pro45.0
19Anthropic logoClaude Opus 4.7 (Fast)44.4
20Alibaba Qwen logoQwen3.5-35B-A3B44.1
21z-ai logoGLM 5.243.1
22arcee-ai logoTrinity Large Thinking42.6
23Alibaba Qwen logoQwen3 Coder Next42.1
24Anthropic logoClaude Sonnet 4.640.8
25inception logoMercury 239.7
26Google DeepMind logoGemini 3.5 Flash37.5
27Alibaba Qwen logoQwen3.5-9B37.4
28DeepSeek logoDeepSeek V4 Pro36.4
29OpenAI logoo336.1
30xAI logoGrok Code Fast 135.6
31minimax logoMiniMax M335.4
32Google DeepMind logoGemini 2.5 Pro32.7
33Alibaba logoQwen3.5 4B32.5
34DeepSeek logoDeepSeek V4 Flash31.1
35Alibaba Qwen logoQwen3.7 Max30.6
36moonshotai logoKimi K2.630.3
37OpenAI logoGPT-5.4 Mini30.2
38z-ai logoGLM 5.129.9
39moonshotai logoKimi K2.7 Code29.6
40xiaomi logoMiMo-V2.5-Pro29.1
41
U
Muse Spark
28.7
42Alibaba Qwen logoQwen3.6 Plus27.6
43OpenAI logoGPT-5.4 Nano27.5
44NVIDIA logoNemotron 3 Ultra27.4
45Alibaba Qwen logoQwen3.6 27B27.0
46Google DeepMind logoGemini 3.1 Flash Lite Preview25.7
47minimax logoMiniMax M2.725.6
48Mistral AI logoMistral Medium 3.125.3
49xAI logoGrok 4.324.1
50Alibaba Qwen logoQwen3 Next 80B A3B Instruct (free)23.6
51Mistral AI logoMistral Small 423.4
52Alibaba logoQwen3.5 2B23.0
53stepfun logoStep 3.7 Flash21.5
54Alibaba Qwen logoQwen3.6 35B A3B21.4
55Google DeepMind logoGemini 3.1 Pro Preview21.4
56DeepSeek logoR1 052820.8
57Alibaba Qwen logoQwen3.7 Plus20.8
58Alibaba Qwen logoQwen3.5-122B-A10B20.7
59Alibaba Qwen logoQwen3.5 397B A17B19.9
60prime-intellect logoINTELLECT-319.8
61Mistral AI logoMistral Medium 3.519.0
62
I
Ring-2.6-1T
18.9
63Alibaba Qwen logoQwen3 Coder 480B A35B (free)18.3
64Alibaba logoQwen3.5 0.8B15.9
65Google DeepMind logoGemma 4 31B (free)14.4
66Alibaba Qwen logoQwen3 Next 80B A3B Instruct14.2
67OpenAI logogpt-oss-120b (free)13.2
68Google DeepMind logoGemini 2.5 Flash Lite11.7
69Google DeepMind logoGemma 4 26B A4B (free)11.0
70NVIDIA logoNVIDIA Nemotron Nano 9B V29.4
71
N
Nanbeige4.1 3B
7.2
72liquid logoLFM2.5-1.2B-Thinking (free)6.5
73Google DeepMind logoGemini 3.1 Flash Lite6.2
74Cohere logoCommand A5.1
75ibm-granite logoGranite 4.0 Micro4.2
76NVIDIA logoLlama 3.1 Nemotron Ultra 253B v13.8
77liquid logoLFM2-24B-A2B3.7
78liquid logoLFM2.5-1.2B-Instruct (free)3.6
79OpenAI logogpt-oss-20b (free)3.1
80Microsoft logoPhi 4 Mini Instruct2.7
81upstage logoSolar Pro 32.7
82ibm-granite logoGranite 4.1 8B1.3
83Meta logoLlama 4 Maverick1.3
84Meta logoLlama 4 Scout1.1
Details
Category
Knowledge
Creator
Artificial Analysis
Max score
60
Modality
Text
Scoring
Weighted composite of multiple agentic benchmark scores. Components include SWE-bench, tool use, planning, and error recovery. Higher is better.
Models
84
Updated
2026-06-16
Tests
Multi-step tool usePlanningError recoveryAutonomous task completion
Does not test
VisionLong contextKnowledge recallCreative writing
Gecko's Take

The Agentic Index is the single number that matters most for 2026. If you are building agents, this is your shortlist filter. If you are investing, this predicts which provider captures the agent platform market.

Same category · related evaluations