Compare · ModelsLive · 2 picked · head to head

Gemini 3.6 Flash vs GPT-5

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Gemini 3.6 Flash wins 13 of 14 shared benchmarks. Leads in agentic · reasoning · knowledge.

Category leads
agentic·Gemini 3.6 Flashreasoning·Gemini 3.6 Flashknowledge·Gemini 3.6 Flashgeneral·Gemini 3.6 Flashmath·Gemini 3.6 Flashcoding·GPT-5
Hype vs Reality
Gemini 3.6 Flash
#91 by perf·#5 by attention
DESERVED
GPT-5
#100 by perf·#4 by attention
DESERVED
Best value
2.5x better value than GPT-5
Gemini 3.6 Flash
23.9 pts/$
$2.25/M
GPT-5
9.4 pts/$
$5.63/M
Vendor risk
Google DeepMind logo
Google DeepMind
$4.20T·Tier 1
Low risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Head to head
Gemini 3.6 FlashGPT-5
APEX-Agents
Gemini 3.6 Flash leads by +28.6
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments.
Gemini 3.6 Flash
46.9
GPT-5
18.3
ARC-AGI
Gemini 3.6 Flash leads by +25.5
ARC-AGI · the original Abstraction and Reasoning Corpus, testing whether AI can solve novel visual pattern recognition tasks without memorization.
Gemini 3.6 Flash
91.2
GPT-5
65.7
ARC-AGI-2
Gemini 3.6 Flash leads by +50.6
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
Gemini 3.6 Flash
60.4
GPT-5
9.9
Chess Puzzles
Gemini 3.6 Flash leads by +6.3
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Gemini 3.6 Flash
40.0
GPT-5
33.7
Dtbench
Gemini 3.6 Flash leads by +8.0
Gemini 3.6 Flash
92.5
GPT-5
84.5
FrontierMath-Tier-4-v2-Private
Gemini 3.6 Flash
21.9
GPT-5
21.9
FrontierMath-Tiers-1-3-v2-Private
Gemini 3.6 Flash leads by +3.5
Gemini 3.6 Flash
59.0
GPT-5
55.4
GPQA diamond
Gemini 3.6 Flash leads by +10.6
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Gemini 3.6 Flash
92.2
GPT-5
81.6
Lmca
Gemini 3.6 Flash leads by +5.8
Gemini 3.6 Flash
52.9
GPT-5
47.0
Mystery Game Puzzles
Gemini 3.6 Flash leads by +7.7
Gemini 3.6 Flash
22.9
GPT-5
15.2
OTIS Mock AIME 2024-2025
Gemini 3.6 Flash leads by +2.8
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Gemini 3.6 Flash
94.2
GPT-5
91.4
Proofbench
Gemini 3.6 Flash leads by +18.0
Gemini 3.6 Flash
36.0
GPT-5
18.0
SimpleQA Verified
Gemini 3.6 Flash leads by +16.1
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Gemini 3.6 Flash
66.2
GPT-5
50.1
WeirdML
GPT-5 leads by +4.6
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Gemini 3.6 Flash
56.1
GPT-5
60.7
Full benchmark table
BenchmarkGemini 3.6 FlashGPT-5
APEX-Agents
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments.
46.918.3
ARC-AGI
ARC-AGI · the original Abstraction and Reasoning Corpus, testing whether AI can solve novel visual pattern recognition tasks without memorization.
91.265.7
ARC-AGI-2
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
60.49.9
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
40.033.7
Dtbench
92.584.5
FrontierMath-Tier-4-v2-Private
21.921.9
FrontierMath-Tiers-1-3-v2-Private
59.055.4
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
92.281.6
Lmca
52.947.0
Mystery Game Puzzles
22.915.2
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
94.291.4
Proofbench
36.018.0
SimpleQA Verified
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
66.250.1
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
56.160.7
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Google DeepMind logoGemini 3.6 Flash$0.75$3.751.0M tokens (~524 books)$15.00
OpenAI logoGPT-5$1.25$10.00400K tokens (~200 books)$34.38