Compare · ModelsLive · 2 picked · head to head
gpt-oss-20b vs Grok 4
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Grok 4 wins on 10/10 benchmarks
Grok 4 wins 10 of 10 shared benchmarks. Leads in knowledge · language · math.
Category leads
knowledge·Grok 4language·Grok 4math·Grok 4reasoning·Grok 4coding·Grok 4
Hype vs Reality
Attention vs performance
gpt-oss-20b
#97 by perf·no signal
Grok 4
#87 by perf·#15 by attention
Vendor risk
Who is behind the model
OpenAI
$840.0B·Tier 1
xAI
$250.0B·Tier 1
Head to head
10 benchmarks · 2 models
gpt-oss-20bGrok 4
Chess Puzzles
Grok 4 leads by +24.2
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
gpt-oss-20b
0.0
Grok 4
24.2
GPQA diamond
Grok 4 leads by +34.9
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
gpt-oss-20b
47.7
Grok 4
82.7
HELM · GPQA
Grok 4 leads by +13.2
gpt-oss-20b
59.4
Grok 4
72.6
HELM · IFEval
Grok 4 leads by +21.7
gpt-oss-20b
73.2
Grok 4
94.9
HELM · MMLU-Pro
Grok 4 leads by +11.1
gpt-oss-20b
74.0
Grok 4
85.1
HELM · Omni-MATH
Grok 4 leads by +3.8
gpt-oss-20b
56.5
Grok 4
60.3
HELM · WildBench
Grok 4 leads by +6.0
gpt-oss-20b
73.7
Grok 4
79.7
OTIS Mock AIME 2024-2025
Grok 4 leads by +18.7
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
gpt-oss-20b
65.2
Grok 4
84.0
Terminal Bench
Grok 4 leads by +23.8
Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence.
gpt-oss-20b
3.4
Grok 4
27.2
WeirdML
Grok 4 leads by +4.8
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
gpt-oss-20b
40.9
Grok 4
45.7
Full benchmark table
| Benchmark | gpt-oss-20b | Grok 4 |
|---|---|---|
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 0.0 | 24.2 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 47.7 | 82.7 |
HELM · GPQA | 59.4 | 72.6 |
HELM · IFEval | 73.2 | 94.9 |
HELM · MMLU-Pro | 74.0 | 85.1 |
HELM · Omni-MATH | 56.5 | 60.3 |
HELM · WildBench | 73.7 | 79.7 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 65.2 | 84.0 |
Terminal Bench Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence. | 3.4 | 27.2 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 40.9 | 45.7 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.02 | $0.09 | 131K tokens (~66 books) | $0.36 | |
| — | — | — | — |
People also compared