Home/Models/Grok 4.20
xAI logo

Grok 4.20

by xAI · Released Mar 2026

Multimodal2M Context
70.8
avg score
Rank #48
Compare
Better than 82% of all models
Context
2.0M tokens (~1,000 books)
Input $/1M
$1.25
Output $/1M
$2.50
Type
multimodal
License
Proprietary
Benchmarks
4 tested
Data updated today
About

Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

Tested on 4 benchmarks with 66.1% average. Top scores: ARC-AGI (89.5%), ARC-AGI-2 (65.1%), Terminal Bench (57.3%).

Looking for similar performance at lower cost?
Qwen3.6 Plus scores 69.8 (99% as good) at $0.33/1M input · 74% cheaper
Capabilities
coding
54.8
#72 globally
reasoning
77.3
#18 globally
Benchmark Scores
Compare All
Tested on 4 benchmarks · Ranked across 2 categories
Score Distribution (all 274 models)
0255075100
▲ You are here
Terminal Bench

Complex terminal-based engineering tasks. Models must use command-line tools, navigate filesystems, and debug systems through shell interaction.

57.3
WeirdML

Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.

52.3
ARC-AGI

Abstraction and Reasoning Corpus. Tests fluid intelligence through novel visual pattern recognition puzzles. Core measure of general intelligence.

89.5
ARC-AGI-2

ARC-AGI 2, harder sequel to ARC. More complex abstract reasoning patterns that test generalization ability beyond training data.

65.1
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Links
Documentation
Community
BenchGecko API
grok-4-20
Specifications
  • Typemultimodal
  • Context2.0M tokens (~1,000 books)
  • ReleasedMar 2026
  • LicenseProprietary
  • StatusActive
  • Cost / Message~$0.005
Available On
xAI logoxAI$1.25
Categories
Share & Export
Tweet
Grok 4.20 is a proprietary multimodal AI model by xAI, released in March 2026. It has an average benchmark score of 70.8. Context window: 2M tokens.