Compare · ModelsLive · 2 picked · head to head
Qwen2.5 72B Instruct vs Qwen2.5 Coder 14B Instruct
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Qwen2.5 72B Instruct wins on 10/11 benchmarks
Qwen2.5 72B Instruct wins 10 of 11 shared benchmarks. Leads in knowledge · general · language.
Category leads
coding·Qwen2.5 Coder 14B Instructknowledge·Qwen2.5 72B Instructgeneral·Qwen2.5 72B Instructlanguage·Qwen2.5 72B Instructmath·Qwen2.5 72B Instructreasoning·Qwen2.5 72B Instruct
Hype vs Reality
Attention vs performance
Qwen2.5 72B Instruct
#107 by perf·no signal
Qwen2.5 Coder 14B Instruct
#120 by perf·no signal
Best value
Qwen2.5 72B Instruct
Qwen2.5 72B Instruct
135.8 pts/$
$0.38/M
Qwen2.5 Coder 14B Instruct
—
no price
Vendor risk
Who is behind the model
Alibaba (Qwen)
$293.0B·Tier 1
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
11 benchmarks · 2 models
Qwen2.5 72B InstructQwen2.5 Coder 14B Instruct
Aider · Code Editing
Qwen2.5 Coder 14B Instruct leads by +3.8
Qwen2.5 72B Instruct
65.4
Qwen2.5 Coder 14B Instruct
69.2
ARC AI2
Qwen2.5 72B Instruct leads by +38.0
AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval.
Qwen2.5 72B Instruct
92.7
Qwen2.5 Coder 14B Instruct
54.7
HellaSwag
Qwen2.5 72B Instruct leads by +6.1
HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.
Qwen2.5 72B Instruct
79.7
Qwen2.5 Coder 14B Instruct
73.6
BBH (HuggingFace)
Qwen2.5 72B Instruct leads by +17.6
Qwen2.5 72B Instruct
61.9
Qwen2.5 Coder 14B Instruct
44.2
GPQA
Qwen2.5 72B Instruct leads by +9.4
Qwen2.5 72B Instruct
16.7
Qwen2.5 Coder 14B Instruct
7.3
IFEval
Qwen2.5 72B Instruct leads by +17.3
Qwen2.5 72B Instruct
86.4
Qwen2.5 Coder 14B Instruct
69.1
MATH Level 5
Qwen2.5 72B Instruct leads by +27.3
Qwen2.5 72B Instruct
59.8
Qwen2.5 Coder 14B Instruct
32.5
MMLU-PRO
Qwen2.5 72B Instruct leads by +18.7
Qwen2.5 72B Instruct
51.4
Qwen2.5 Coder 14B Instruct
32.7
MUSR
Qwen2.5 72B Instruct leads by +4.7
Qwen2.5 72B Instruct
11.7
Qwen2.5 Coder 14B Instruct
7.0
MMLU
Qwen2.5 72B Instruct leads by +13.5
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Qwen2.5 72B Instruct
80.4
Qwen2.5 Coder 14B Instruct
66.9
Winogrande
Qwen2.5 72B Instruct leads by +11.0
WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs.
Qwen2.5 72B Instruct
64.6
Qwen2.5 Coder 14B Instruct
53.6
Full benchmark table
| Benchmark | Qwen2.5 72B Instruct | Qwen2.5 Coder 14B Instruct |
|---|---|---|
Aider · Code Editing | 65.4 | 69.2 |
ARC AI2 AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval. | 92.7 | 54.7 |
HellaSwag HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios. | 79.7 | 73.6 |
BBH (HuggingFace) | 61.9 | 44.2 |
GPQA | 16.7 | 7.3 |
IFEval | 86.4 | 69.1 |
MATH Level 5 | 59.8 | 32.5 |
MMLU-PRO | 51.4 | 32.7 |
MUSR | 11.7 | 7.0 |
MMLU Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge. | 80.4 | 66.9 |
Winogrande WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs. | 64.6 | 53.6 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.36 | $0.40 | 131K tokens (~66 books) | $3.70 | |
| — | — | — | — |