Compare · ModelsLive · 2 picked · head to head
Claude 3.5 Haiku vs GPT-5.4 Mini
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
GPT-5.4 Mini wins on 6/6 benchmarks
GPT-5.4 Mini wins 6 of 6 shared benchmarks. Leads in general · math · knowledge.
Category leads
general·GPT-5.4 Minimath·GPT-5.4 Miniknowledge·GPT-5.4 Minicoding·GPT-5.4 Mini
Hype vs Reality
Attention vs performance
Claude 3.5 Haiku
#210 by perf·no signal
GPT-5.4 Mini
#209 by perf·#4 by attention
Vendor risk
Who is behind the model
Anthropic
$965.0B·Tier 1
OpenAI
$840.0B·Tier 1
Head to head
6 benchmarks · 2 models
Claude 3.5 HaikuGPT-5.4 Mini
Dtbench
GPT-5.4 Mini leads by +38.9
Claude 3.5 Haiku
27.8
GPT-5.4 Mini
66.7
FrontierMath-2025-02-28-Private
GPT-5.4 Mini leads by +27.7
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Claude 3.5 Haiku
0.6
GPT-5.4 Mini
28.3
GPQA diamond
GPT-5.4 Mini leads by +65.0
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Claude 3.5 Haiku
17.5
GPT-5.4 Mini
82.5
OTIS Mock AIME 2024-2025
GPT-5.4 Mini leads by +84.7
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Claude 3.5 Haiku
4.2
GPT-5.4 Mini
88.9
SimpleQA Verified
GPT-5.4 Mini leads by +22.7
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Claude 3.5 Haiku
6.7
GPT-5.4 Mini
29.4
WeirdML
GPT-5.4 Mini leads by +29.6
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Claude 3.5 Haiku
30.7
GPT-5.4 Mini
60.3
Full benchmark table
| Benchmark | Claude 3.5 Haiku | GPT-5.4 Mini |
|---|---|---|
Dtbench | 27.8 | 66.7 |
FrontierMath-2025-02-28-Private FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning. | 0.6 | 28.3 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 17.5 | 82.5 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 4.2 | 88.9 |
SimpleQA Verified SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information. | 6.7 | 29.4 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 30.7 | 60.3 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| — | — | — | — | |
| $0.75 | $4.50 | 400K tokens (~200 books) | $16.88 |
People also compared