Compare · ModelsLive · 2 picked · head to head
gpt-oss-20b vs o3 Mini
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
o3 Mini wins on 7/7 benchmarks
o3 Mini wins 7 of 7 shared benchmarks. Leads in arena · knowledge · general.
Category leads
arena·o3 Miniknowledge·o3 Minigeneral·o3 Minimath·o3 Minicoding·o3 Mini
Hype vs Reality
Attention vs performance
gpt-oss-20b
#97 by perf·no signal
o3 Mini
#239 by perf·no signal
Best value
gpt-oss-20b
83.2x better value than o3 Mini
gpt-oss-20b
983.3 pts/$
$0.05/M
o3 Mini
11.8 pts/$
$2.75/M
Vendor risk
Who is behind the model
OpenAI
$840.0B·Tier 1
OpenAI
$840.0B·Tier 1
Head to head
7 benchmarks · 2 models
gpt-oss-20bo3 Mini
Chatbot Arena Elo · Overall
o3 Mini leads by +30.5
gpt-oss-20b
1317.6
o3 Mini
1348.1
Chess Puzzles
o3 Mini leads by +12.7
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
gpt-oss-20b
0.0
o3 Mini
12.7
Dtbench
o3 Mini leads by +1.3
gpt-oss-20b
46.7
o3 Mini
48.0
GPQA diamond
o3 Mini leads by +21.6
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
gpt-oss-20b
47.7
o3 Mini
69.4
Lmca
o3 Mini leads by +5.3
gpt-oss-20b
17.1
o3 Mini
22.3
OTIS Mock AIME 2024-2025
o3 Mini leads by +11.7
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
gpt-oss-20b
65.2
o3 Mini
76.9
WeirdML
o3 Mini leads by +2.8
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
gpt-oss-20b
40.9
o3 Mini
43.7
Full benchmark table
| Benchmark | gpt-oss-20b | o3 Mini |
|---|---|---|
Chatbot Arena Elo · Overall | 1317.6 | 1348.1 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 0.0 | 12.7 |
Dtbench | 46.7 | 48.0 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 47.7 | 69.4 |
Lmca | 17.1 | 22.3 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 65.2 | 76.9 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 40.9 | 43.7 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.02 | $0.09 | 131K tokens (~66 books) | $0.36 | |
| $1.10 | $4.40 | 200K tokens (~100 books) | $19.25 |
People also compared