Compare · ModelsLive · 2 picked · head to head

Claude Fable 5.1 vs Gemini 3.1 Pro Preview

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Claude Fable 5.1 wins 20 of 24 shared benchmarks. Leads in speed · agentic · general.

Category leads
speed·Claude Fable 5.1agentic·Claude Fable 5.1reasoning·Gemini 3.1 Pro Previewknowledge·Gemini 3.1 Pro Previewgeneral·Claude Fable 5.1math·Claude Fable 5.1coding·Claude Fable 5.1
Hype vs Reality
Claude Fable 5.1
#16 by perf·#12 by attention
UNDERRATED
Gemini 3.1 Pro Preview
#137 by perf·#5 by attention
DESERVED
Best value
2.9x better value than Claude Fable 5.1
Claude Fable 5.1
2.4 pts/$
$30.00/M
Gemini 3.1 Pro Preview
6.9 pts/$
$7.00/M
Vendor risk
Anthropic logo
Anthropic
$965.0B·Tier 1
Medium risk
Google DeepMind logo
Google DeepMind
$4.20T·Tier 1
Low risk
Head to head
Claude Fable 5.1Gemini 3.1 Pro Preview
Artificial Analysis · CritPt
Claude Fable 5.1 leads by +12.0
Claude Fable 5.1
29.7
Gemini 3.1 Pro Preview
17.7
Artificial Analysis · GDPval
Claude Fable 5.1 leads by +47.9
Claude Fable 5.1
61.7
Gemini 3.1 Pro Preview
13.8
Artificial Analysis · GPQA Diamond
Gemini 3.1 Pro Preview leads by +0.4
Claude Fable 5.1
93.7
Gemini 3.1 Pro Preview
94.1
Artificial Analysis · Humanity's Last Exam
Claude Fable 5.1 leads by +12.1
Claude Fable 5.1
59.1
Gemini 3.1 Pro Preview
47.0
Artificial Analysis · Long Context Reasoning
Claude Fable 5.1 leads by +3.3
Claude Fable 5.1
85.3
Gemini 3.1 Pro Preview
82.0
Artificial Analysis · Quality Index
Claude Fable 5.1 leads by +23.6
Claude Fable 5.1
53.4
Gemini 3.1 Pro Preview
29.7
Artificial Analysis · SciCode
Claude Fable 5.1 leads by +4.4
Claude Fable 5.1
63.1
Gemini 3.1 Pro Preview
58.7
APEX-Agents
Claude Fable 5.1 leads by +33.3
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments.
Claude Fable 5.1
68.6
Gemini 3.1 Pro Preview
35.3
ARC-AGI
Gemini 3.1 Pro Preview leads by +0.5
ARC-AGI · the original Abstraction and Reasoning Corpus, testing whether AI can solve novel visual pattern recognition tasks without memorization.
Claude Fable 5.1
97.5
Gemini 3.1 Pro Preview
98.0
ARC-AGI-2
Claude Fable 5.1 leads by +12.9
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
Claude Fable 5.1
90.0
Gemini 3.1 Pro Preview
77.1
Chess Puzzles
Gemini 3.1 Pro Preview leads by +8.4
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Claude Fable 5.1
44.2
Gemini 3.1 Pro Preview
52.6
Dtbench
Claude Fable 5.1 leads by +0.9
Claude Fable 5.1
96.0
Gemini 3.1 Pro Preview
95.1
Ebr Bench
Claude Fable 5.1 leads by +42.9
Claude Fable 5.1
57.1
Gemini 3.1 Pro Preview
14.3
FrontierMath-Tier-4-v2-Private
Claude Fable 5.1 leads by +61.0
Claude Fable 5.1
87.8
Gemini 3.1 Pro Preview
26.8
FrontierMath-Tiers-1-3-v2-Private
Claude Fable 5.1 leads by +30.5
Claude Fable 5.1
90.2
Gemini 3.1 Pro Preview
59.6
Furniture Assembly
Claude Fable 5.1 leads by +57.1
Claude Fable 5.1
57.1
Gemini 3.1 Pro Preview
0.0
HLE
Claude Fable 5.1 leads by +0.1
HLE (Humanity's Last Exam) · a reasoning benchmark designed to be the hardest public evaluation of AI. Questions span mathematics, physics, philosophy, and logic · curated to be at or beyond the frontier of human expert capability. Tested with and without tool augmentation. Claude Opus 4.7 scores 46.9% without tools and 54.7% with tools · making it one of the few benchmarks where the top score is below 60%.
Claude Fable 5.1
43.8
Gemini 3.1 Pro Preview
43.7
Lmca
Claude Fable 5.1 leads by +13.7
Claude Fable 5.1
77.0
Gemini 3.1 Pro Preview
63.3
Mirrorcode
Claude Fable 5.1 leads by +64.4
Claude Fable 5.1
73.3
Gemini 3.1 Pro Preview
8.9
Mystery Game Puzzles
Claude Fable 5.1 leads by +26.4
Claude Fable 5.1
53.7
Gemini 3.1 Pro Preview
27.3
OTIS Mock AIME 2024-2025
Claude Fable 5.1 leads by +4.4
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Claude Fable 5.1
100.0
Gemini 3.1 Pro Preview
95.6
Proofbench
Claude Fable 5.1 leads by +74.0
Claude Fable 5.1
100.0
Gemini 3.1 Pro Preview
26.0
SimpleQA Verified
Gemini 3.1 Pro Preview leads by +2.7
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Claude Fable 5.1
70.8
Gemini 3.1 Pro Preview
73.5
WeirdML
Claude Fable 5.1 leads by +20.8
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Claude Fable 5.1
92.9
Gemini 3.1 Pro Preview
72.1
Full benchmark table
BenchmarkClaude Fable 5.1Gemini 3.1 Pro Preview
Artificial Analysis · CritPt
29.717.7
Artificial Analysis · GDPval
61.713.8
Artificial Analysis · GPQA Diamond
93.794.1
Artificial Analysis · Humanity's Last Exam
59.147.0
Artificial Analysis · Long Context Reasoning
85.382.0
Artificial Analysis · Quality Index
53.429.7
Artificial Analysis · SciCode
63.158.7
APEX-Agents
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments.
68.635.3
ARC-AGI
ARC-AGI · the original Abstraction and Reasoning Corpus, testing whether AI can solve novel visual pattern recognition tasks without memorization.
97.598.0
ARC-AGI-2
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
90.077.1
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
44.252.6
Dtbench
96.095.1
Ebr Bench
57.114.3
FrontierMath-Tier-4-v2-Private
87.826.8
FrontierMath-Tiers-1-3-v2-Private
90.259.6
Furniture Assembly
57.10.0
HLE
HLE (Humanity's Last Exam) · a reasoning benchmark designed to be the hardest public evaluation of AI. Questions span mathematics, physics, philosophy, and logic · curated to be at or beyond the frontier of human expert capability. Tested with and without tool augmentation. Claude Opus 4.7 scores 46.9% without tools and 54.7% with tools · making it one of the few benchmarks where the top score is below 60%.
43.843.7
Lmca
77.063.3
Mirrorcode
73.38.9
Mystery Game Puzzles
53.727.3
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
100.095.6
Proofbench
100.026.0
SimpleQA Verified
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
70.873.5
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
92.972.1
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Anthropic logoClaude Fable 5.1$10.00$50.001.0M tokens (~500 books)$200.00
Google DeepMind logoGemini 3.1 Pro Preview$2.00$12.001.0M tokens (~524 books)$45.00