Compare · ModelsLive · 2 picked · head to head
Gemini 2.5 Pro vs Grok 4 Fast
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Grok 4 Fast wins on 4/6 benchmarks
Grok 4 Fast wins 4 of 6 shared benchmarks. Leads in reasoning · general · knowledge.
Category leads
reasoning·Grok 4 Fastgeneral·Grok 4 Fastknowledge·Grok 4 Fastcoding·Gemini 2.5 Pro
Hype vs Reality
Attention vs performance
Gemini 2.5 Pro
#116 by perf·no signal
Grok 4 Fast
#96 by perf·#15 by attention
Vendor risk
Who is behind the model
Google DeepMind
$4.20T·Tier 1
xAI
$250.0B·Tier 1
Head to head
6 benchmarks · 2 models
Gemini 2.5 ProGrok 4 Fast
ARC-AGI
Grok 4 Fast leads by +7.5
ARC-AGI · the original Abstraction and Reasoning Corpus, testing whether AI can solve novel visual pattern recognition tasks without memorization.
Gemini 2.5 Pro
41.0
Grok 4 Fast
48.5
ARC-AGI-2
Grok 4 Fast leads by +0.4
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
Gemini 2.5 Pro
4.9
Grok 4 Fast
5.3
Dtbench
Grok 4 Fast leads by +0.5
Gemini 2.5 Pro
70.7
Grok 4 Fast
71.1
Fiction.LiveBench
Grok 4 Fast leads by +2.7
Fiction.LiveBench · a continuously updated benchmark using recently published fiction to test reading comprehension and reasoning, preventing data contamination.
Gemini 2.5 Pro
91.7
Grok 4 Fast
94.4
Lech Mazur Writing
Gemini 2.5 Pro leads by +2.7
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
Gemini 2.5 Pro
83.8
Grok 4 Fast
81.1
WeirdML
Gemini 2.5 Pro leads by +11.2
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Gemini 2.5 Pro
54.0
Grok 4 Fast
42.9
Full benchmark table
| Benchmark | Gemini 2.5 Pro | Grok 4 Fast |
|---|---|---|
ARC-AGI ARC-AGI · the original Abstraction and Reasoning Corpus, testing whether AI can solve novel visual pattern recognition tasks without memorization. | 41.0 | 48.5 |
ARC-AGI-2 ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data. | 4.9 | 5.3 |
Dtbench | 70.7 | 71.1 |
Fiction.LiveBench Fiction.LiveBench · a continuously updated benchmark using recently published fiction to test reading comprehension and reasoning, preventing data contamination. | 91.7 | 94.4 |
Lech Mazur Writing Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication. | 83.8 | 81.1 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 54.0 | 42.9 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $1.25 | $10.00 | 1.0M tokens (~524 books) | $34.38 | |
| — | — | — | — |