Compare · ModelsLive · 2 picked · head to head
Llama 3.1 8B Instruct vs o3 Mini
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
o3 Mini wins on 8/8 benchmarks
o3 Mini wins 8 of 8 shared benchmarks. Leads in arena · knowledge · general.
Category leads
arena·o3 Miniknowledge·o3 Minigeneral·o3 Minimath·o3 Minicoding·o3 Mini
Hype vs Reality
Attention vs performance
Llama 3.1 8B Instruct
#275 by perf·#18 by attention
o3 Mini
#239 by perf·no signal
Best value
Llama 3.1 8B Instruct
81.9x better value than o3 Mini
Llama 3.1 8B Instruct
968.0 pts/$
$0.03/M
o3 Mini
11.8 pts/$
$2.75/M
Vendor risk
Who is behind the model
Meta AI
$1.87T·Tier 1
OpenAI
$840.0B·Tier 1
Head to head
8 benchmarks · 2 models
Llama 3.1 8B Instructo3 Mini
Chatbot Arena Elo · Overall
o3 Mini leads by +136.8
Llama 3.1 8B Instruct
1211.2
o3 Mini
1348.1
Chess Puzzles
o3 Mini leads by +12.7
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Llama 3.1 8B Instruct
0.0
o3 Mini
12.7
Dtbench
o3 Mini leads by +29.8
Llama 3.1 8B Instruct
18.2
o3 Mini
48.0
GPQA diamond
o3 Mini leads by +66.8
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Llama 3.1 8B Instruct
2.6
o3 Mini
69.4
Lmca
o3 Mini leads by +16.0
Llama 3.1 8B Instruct
6.3
o3 Mini
22.3
MATH level 5
o3 Mini leads by +73.6
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Llama 3.1 8B Instruct
22.9
o3 Mini
96.5
OTIS Mock AIME 2024-2025
o3 Mini leads by +75.4
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Llama 3.1 8B Instruct
1.6
o3 Mini
76.9
WeirdML
o3 Mini leads by +42.0
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Llama 3.1 8B Instruct
1.7
o3 Mini
43.7
Full benchmark table
| Benchmark | Llama 3.1 8B Instruct | o3 Mini |
|---|---|---|
Chatbot Arena Elo · Overall | 1211.2 | 1348.1 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 0.0 | 12.7 |
Dtbench | 18.2 | 48.0 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 2.6 | 69.4 |
Lmca | 6.3 | 22.3 |
MATH level 5 MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics. | 22.9 | 96.5 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 1.6 | 76.9 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 1.7 | 43.7 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.02 | $0.03 | 131K tokens (~66 books) | $0.23 | |
| $1.10 | $4.40 | 200K tokens (~100 books) | $19.25 |
People also compared