Compare · ModelsLive · 2 picked · head to head

GPT-5.4 Mini vs Claude 3.5 Haiku

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

GPT-5.4 Mini wins 5 of 5 shared benchmarks. Leads in math · knowledge · coding.

Category leads
math·GPT-5.4 Miniknowledge·GPT-5.4 Minicoding·GPT-5.4 Mini
Hype vs Reality
GPT-5.4 Mini
#182 by perf·no signal
QUIET
Claude 3.5 Haiku
#187 by perf·no signal
QUIET
Best value
GPT-5.4 Mini
14.6 pts/$
$2.63/M
Claude 3.5 Haiku
no price
Vendor risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Anthropic logo
Anthropic
$380.0B·Tier 1
Medium risk
Head to head
GPT-5.4 MiniClaude 3.5 Haiku
FrontierMath-2025-02-28-Private
GPT-5.4 Mini leads by +27.9
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
GPT-5.4 Mini
28.3
Claude 3.5 Haiku
0.3
GPQA diamond
GPT-5.4 Mini leads by +60.6
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
GPT-5.4 Mini
78.1
Claude 3.5 Haiku
17.5
OTIS Mock AIME 2024-2025
GPT-5.4 Mini leads by +83.0
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
GPT-5.4 Mini
87.2
Claude 3.5 Haiku
4.2
SimpleQA Verified
GPT-5.4 Mini leads by +21.9
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
GPT-5.4 Mini
28.6
Claude 3.5 Haiku
6.7
WeirdML
GPT-5.4 Mini leads by +29.6
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
GPT-5.4 Mini
60.3
Claude 3.5 Haiku
30.7
Full benchmark table
BenchmarkGPT-5.4 MiniClaude 3.5 Haiku
FrontierMath-2025-02-28-Private
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
28.30.3
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
78.117.5
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
87.24.2
SimpleQA Verified
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
28.66.7
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
60.330.7
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
OpenAI logoGPT-5.4 Mini$0.75$4.50400K tokens (~200 books)$16.88
Anthropic logoClaude 3.5 Haiku