Compare · ModelsLive · 2 picked · head to head
Claude Haiku 4.5 vs Qwen2.5 7B Instruct
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Claude Haiku 4.5 wins on 6/6 benchmarks
Claude Haiku 4.5 wins 6 of 6 shared benchmarks. Leads in knowledge · general · math.
Category leads
knowledge·Claude Haiku 4.5general·Claude Haiku 4.5math·Claude Haiku 4.5
Hype vs Reality
Attention vs performance
Claude Haiku 4.5
#231 by perf·#13 by attention
Qwen2.5 7B Instruct
#273 by perf·#2 by attention
Best value
Qwen2.5 7B Instruct
14.5x better value than Claude Haiku 4.5
Claude Haiku 4.5
11.3 pts/$
$3.00/M
Qwen2.5 7B Instruct
164.0 pts/$
$0.15/M
Vendor risk
Who is behind the model
Anthropic
$965.0B·Tier 1
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
6 benchmarks · 2 models
Claude Haiku 4.5Qwen2.5 7B Instruct
Balrog
Claude Haiku 4.5 leads by +23.4
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
Claude Haiku 4.5
31.2
Qwen2.5 7B Instruct
7.8
Chess Puzzles
Claude Haiku 4.5 leads by +3.2
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Claude Haiku 4.5
3.2
Qwen2.5 7B Instruct
0.0
Dtbench
Claude Haiku 4.5 leads by +43.1
Claude Haiku 4.5
56.0
Qwen2.5 7B Instruct
12.9
GPQA diamond
Claude Haiku 4.5 leads by +47.6
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Claude Haiku 4.5
61.6
Qwen2.5 7B Instruct
14.0
Lmca
Claude Haiku 4.5 leads by +28.8
Claude Haiku 4.5
36.4
Qwen2.5 7B Instruct
7.5
OTIS Mock AIME 2024-2025
Claude Haiku 4.5 leads by +64.2
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Claude Haiku 4.5
66.6
Qwen2.5 7B Instruct
2.4
Full benchmark table
| Benchmark | Claude Haiku 4.5 | Qwen2.5 7B Instruct |
|---|---|---|
Balrog Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning. | 31.2 | 7.8 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 3.2 | 0.0 |
Dtbench | 56.0 | 12.9 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 61.6 | 14.0 |
Lmca | 36.4 | 7.5 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 66.6 | 2.4 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $1.00 | $5.00 | 200K tokens (~100 books) | $20.00 | |
| $0.10 | $0.20 | 33K tokens (~16 books) | $1.25 |
People also compared