Compare · ModelsLive · 2 picked · head to head
GPT-5.1 vs Grok 4.20
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
GPT-5.1 wins on 7/12 benchmarks
GPT-5.1 wins 7 of 12 shared benchmarks. Leads in knowledge · general.
Category leads
reasoning·Grok 4.20knowledge·GPT-5.1general·GPT-5.1math·Grok 4.20coding·Grok 4.20
Hype vs Reality
Attention vs performance
GPT-5.1
#138 by perf·#4 by attention
Grok 4.20
#129 by perf·#14 by attention
Best value
Grok 4.20
3.0x better value than GPT-5.1
GPT-5.1
8.6 pts/$
$5.63/M
Grok 4.20
26.0 pts/$
$1.88/M
Vendor risk
Who is behind the model
OpenAI
$840.0B·Tier 1
xAI
$250.0B·Tier 1
Head to head
12 benchmarks · 2 models
GPT-5.1Grok 4.20
ARC-AGI
Grok 4.20 leads by +16.7
ARC-AGI · the original Abstraction and Reasoning Corpus, testing whether AI can solve novel visual pattern recognition tasks without memorization.
GPT-5.1
72.8
Grok 4.20
89.5
ARC-AGI-2
Grok 4.20 leads by +47.5
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
GPT-5.1
17.6
Grok 4.20
65.1
Chess Puzzles
GPT-5.1 leads by +8.4
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
GPT-5.1
28.4
Grok 4.20
20.0
Cl Bench
GPT-5.1 leads by +1.5
GPT-5.1
23.7
Grok 4.20
22.2
Cl Bench Life
GPT-5.1 leads by +5.4
GPT-5.1
17.3
Grok 4.20
11.9
Dtbench
GPT-5.1
83.5
Grok 4.20
83.5
GPQA diamond
Grok 4.20 leads by +2.3
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
GPT-5.1
83.5
Grok 4.20
85.8
Lmca
GPT-5.1 leads by +6.2
GPT-5.1
51.6
Grok 4.20
45.5
OTIS Mock AIME 2024-2025
Grok 4.20 leads by +3.6
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
GPT-5.1
88.6
Grok 4.20
92.2
SimpleQA Verified
GPT-5.1 leads by +17.8
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
GPT-5.1
48.0
Grok 4.20
30.2
Terminal Bench
Grok 4.20 leads by +9.7
Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence.
GPT-5.1
47.6
Grok 4.20
57.3
WeirdML
GPT-5.1 leads by +8.5
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
GPT-5.1
60.8
Grok 4.20
52.3
Full benchmark table
| Benchmark | GPT-5.1 | Grok 4.20 |
|---|---|---|
ARC-AGI ARC-AGI · the original Abstraction and Reasoning Corpus, testing whether AI can solve novel visual pattern recognition tasks without memorization. | 72.8 | 89.5 |
ARC-AGI-2 ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data. | 17.6 | 65.1 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 28.4 | 20.0 |
Cl Bench | 23.7 | 22.2 |
Cl Bench Life | 17.3 | 11.9 |
Dtbench | 83.5 | 83.5 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 83.5 | 85.8 |
Lmca | 51.6 | 45.5 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 88.6 | 92.2 |
SimpleQA Verified SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information. | 48.0 | 30.2 |
Terminal Bench Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence. | 47.6 | 57.3 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 60.8 | 52.3 |