Compare · ModelsLive · 2 picked · head to head
Grok 4 vs Qwen2-72B
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Grok 4 wins on 2/2 benchmarks
Grok 4 wins 2 of 2 shared benchmarks. Leads in knowledge · coding.
Category leads
knowledge·Grok 4coding·Grok 4
Hype vs Reality
Attention vs performance
Grok 4
#88 by perf·no signal
Qwen2-72B
#177 by perf·no signal
Vendor risk
Who is behind the model
xAI
$250.0B·Tier 1
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
2 benchmarks · 2 models
Grok 4Qwen2-72B
GPQA diamond
Grok 4 leads by +61.6
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Grok 4
82.7
Qwen2-72B
21.0
WeirdML
Grok 4 leads by +34.4
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Grok 4
45.7
Qwen2-72B
11.3
Full benchmark table
| Benchmark | Grok 4 | Qwen2-72B |
|---|---|---|
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 82.7 | 21.0 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 45.7 | 11.3 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
People also compared