Compare · ModelsLive · 2 picked · head to head
Claude Fable 5.1 vs Gemini 3.1 Pro Preview
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Claude Fable 5.1 wins on 20/24 benchmarks
Claude Fable 5.1 wins 20 of 24 shared benchmarks. Leads in speed · agentic · general.
Category leads
speed·Claude Fable 5.1agentic·Claude Fable 5.1reasoning·Gemini 3.1 Pro Previewknowledge·Gemini 3.1 Pro Previewgeneral·Claude Fable 5.1math·Claude Fable 5.1coding·Claude Fable 5.1
Hype vs Reality
Attention vs performance
Claude Fable 5.1
#16 by perf·#12 by attention
Gemini 3.1 Pro Preview
#137 by perf·#5 by attention
Best value
Gemini 3.1 Pro Preview
2.9x better value than Claude Fable 5.1
Claude Fable 5.1
2.4 pts/$
$30.00/M
Gemini 3.1 Pro Preview
6.9 pts/$
$7.00/M
Vendor risk
Who is behind the model
Anthropic
$965.0B·Tier 1
Google DeepMind
$4.20T·Tier 1
Head to head
24 benchmarks · 2 models
Claude Fable 5.1Gemini 3.1 Pro Preview
Artificial Analysis · CritPt
Claude Fable 5.1 leads by +12.0
Claude Fable 5.1
29.7
Gemini 3.1 Pro Preview
17.7
Artificial Analysis · GDPval
Claude Fable 5.1 leads by +47.9
Claude Fable 5.1
61.7
Gemini 3.1 Pro Preview
13.8
Artificial Analysis · GPQA Diamond
Gemini 3.1 Pro Preview leads by +0.4
Claude Fable 5.1
93.7
Gemini 3.1 Pro Preview
94.1
Artificial Analysis · Humanity's Last Exam
Claude Fable 5.1 leads by +12.1
Claude Fable 5.1
59.1
Gemini 3.1 Pro Preview
47.0
Artificial Analysis · Long Context Reasoning
Claude Fable 5.1 leads by +3.3
Claude Fable 5.1
85.3
Gemini 3.1 Pro Preview
82.0
Artificial Analysis · Quality Index
Claude Fable 5.1 leads by +23.6
Claude Fable 5.1
53.4
Gemini 3.1 Pro Preview
29.7
Artificial Analysis · SciCode
Claude Fable 5.1 leads by +4.4
Claude Fable 5.1
63.1
Gemini 3.1 Pro Preview
58.7
APEX-Agents
Claude Fable 5.1 leads by +33.3
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments.
Claude Fable 5.1
68.6
Gemini 3.1 Pro Preview
35.3
ARC-AGI
Gemini 3.1 Pro Preview leads by +0.5
ARC-AGI · the original Abstraction and Reasoning Corpus, testing whether AI can solve novel visual pattern recognition tasks without memorization.
Claude Fable 5.1
97.5
Gemini 3.1 Pro Preview
98.0
ARC-AGI-2
Claude Fable 5.1 leads by +12.9
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
Claude Fable 5.1
90.0
Gemini 3.1 Pro Preview
77.1
Chess Puzzles
Gemini 3.1 Pro Preview leads by +8.4
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Claude Fable 5.1
44.2
Gemini 3.1 Pro Preview
52.6
Dtbench
Claude Fable 5.1 leads by +0.9
Claude Fable 5.1
96.0
Gemini 3.1 Pro Preview
95.1
Ebr Bench
Claude Fable 5.1 leads by +42.9
Claude Fable 5.1
57.1
Gemini 3.1 Pro Preview
14.3
FrontierMath-Tier-4-v2-Private
Claude Fable 5.1 leads by +61.0
Claude Fable 5.1
87.8
Gemini 3.1 Pro Preview
26.8
FrontierMath-Tiers-1-3-v2-Private
Claude Fable 5.1 leads by +30.5
Claude Fable 5.1
90.2
Gemini 3.1 Pro Preview
59.6
Furniture Assembly
Claude Fable 5.1 leads by +57.1
Claude Fable 5.1
57.1
Gemini 3.1 Pro Preview
0.0
HLE
Claude Fable 5.1 leads by +0.1
HLE (Humanity's Last Exam) · a reasoning benchmark designed to be the hardest public evaluation of AI. Questions span mathematics, physics, philosophy, and logic · curated to be at or beyond the frontier of human expert capability. Tested with and without tool augmentation. Claude Opus 4.7 scores 46.9% without tools and 54.7% with tools · making it one of the few benchmarks where the top score is below 60%.
Claude Fable 5.1
43.8
Gemini 3.1 Pro Preview
43.7
Lmca
Claude Fable 5.1 leads by +13.7
Claude Fable 5.1
77.0
Gemini 3.1 Pro Preview
63.3
Mirrorcode
Claude Fable 5.1 leads by +64.4
Claude Fable 5.1
73.3
Gemini 3.1 Pro Preview
8.9
Mystery Game Puzzles
Claude Fable 5.1 leads by +26.4
Claude Fable 5.1
53.7
Gemini 3.1 Pro Preview
27.3
OTIS Mock AIME 2024-2025
Claude Fable 5.1 leads by +4.4
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Claude Fable 5.1
100.0
Gemini 3.1 Pro Preview
95.6
Proofbench
Claude Fable 5.1 leads by +74.0
Claude Fable 5.1
100.0
Gemini 3.1 Pro Preview
26.0
SimpleQA Verified
Gemini 3.1 Pro Preview leads by +2.7
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Claude Fable 5.1
70.8
Gemini 3.1 Pro Preview
73.5
WeirdML
Claude Fable 5.1 leads by +20.8
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Claude Fable 5.1
92.9
Gemini 3.1 Pro Preview
72.1
Full benchmark table
| Benchmark | Claude Fable 5.1 | Gemini 3.1 Pro Preview |
|---|---|---|
Artificial Analysis · CritPt | 29.7 | 17.7 |
Artificial Analysis · GDPval | 61.7 | 13.8 |
Artificial Analysis · GPQA Diamond | 93.7 | 94.1 |
Artificial Analysis · Humanity's Last Exam | 59.1 | 47.0 |
Artificial Analysis · Long Context Reasoning | 85.3 | 82.0 |
Artificial Analysis · Quality Index | 53.4 | 29.7 |
Artificial Analysis · SciCode | 63.1 | 58.7 |
APEX-Agents APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments. | 68.6 | 35.3 |
ARC-AGI ARC-AGI · the original Abstraction and Reasoning Corpus, testing whether AI can solve novel visual pattern recognition tasks without memorization. | 97.5 | 98.0 |
ARC-AGI-2 ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data. | 90.0 | 77.1 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 44.2 | 52.6 |
Dtbench | 96.0 | 95.1 |
Ebr Bench | 57.1 | 14.3 |
FrontierMath-Tier-4-v2-Private | 87.8 | 26.8 |
FrontierMath-Tiers-1-3-v2-Private | 90.2 | 59.6 |
Furniture Assembly | 57.1 | 0.0 |
HLE HLE (Humanity's Last Exam) · a reasoning benchmark designed to be the hardest public evaluation of AI. Questions span mathematics, physics, philosophy, and logic · curated to be at or beyond the frontier of human expert capability. Tested with and without tool augmentation. Claude Opus 4.7 scores 46.9% without tools and 54.7% with tools · making it one of the few benchmarks where the top score is below 60%. | 43.8 | 43.7 |
Lmca | 77.0 | 63.3 |
Mirrorcode | 73.3 | 8.9 |
Mystery Game Puzzles | 53.7 | 27.3 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 100.0 | 95.6 |
Proofbench | 100.0 | 26.0 |
SimpleQA Verified SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information. | 70.8 | 73.5 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 92.9 | 72.1 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $10.00 | $50.00 | 1.0M tokens (~500 books) | $200.00 | |
| $2.00 | $12.00 | 1.0M tokens (~524 books) | $45.00 |
People also compared