Compare · ModelsLive · 2 picked · head to head
Llama 3.3 70B Instruct (free) vs Gemini 1.5 Pro (May 2024)
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Llama 3.3 70B Instruct (free) wins on 4/7 benchmarks
Llama 3.3 70B Instruct (free) wins 4 of 7 shared benchmarks. Leads in knowledge · math.
Category leads
knowledge·Llama 3.3 70B Instruct (free)math·Llama 3.3 70B Instruct (free)reasoning·Gemini 1.5 Pro (May 2024)coding·Gemini 1.5 Pro (May 2024)
Hype vs Reality
Attention vs performance
Llama 3.3 70B Instruct (free)
#224 by perf·no signal
Gemini 1.5 Pro (May 2024)
#165 by perf·no signal
Best value
Pricing unknown
Llama 3.3 70B Instruct (free)
—
$0.00/M
Gemini 1.5 Pro (May 2024)
—
no price
Vendor risk
Who is behind the model
Meta AI
$1.50T·Tier 1
Google DeepMind
$4.00T·Tier 1
Head to head
7 benchmarks · 2 models
Llama 3.3 70B Instruct (free)Gemini 1.5 Pro (May 2024)
Balrog
Llama 3.3 70B Instruct (free) leads by +2.0
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
Llama 3.3 70B Instruct (free)
23.0
Gemini 1.5 Pro (May 2024)
21.0
GPQA diamond
Llama 3.3 70B Instruct (free) leads by +2.1
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Llama 3.3 70B Instruct (free)
29.9
Gemini 1.5 Pro (May 2024)
27.8
MATH level 5
Llama 3.3 70B Instruct (free) leads by +0.9
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Llama 3.3 70B Instruct (free)
41.6
Gemini 1.5 Pro (May 2024)
40.8
MMLU
Llama 3.3 70B Instruct (free) leads by +0.5
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Llama 3.3 70B Instruct (free)
81.7
Gemini 1.5 Pro (May 2024)
81.2
OTIS Mock AIME 2024-2025
Gemini 1.5 Pro (May 2024) leads by +1.7
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Llama 3.3 70B Instruct (free)
5.0
Gemini 1.5 Pro (May 2024)
6.7
SimpleBench
Gemini 1.5 Pro (May 2024) leads by +8.6
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Llama 3.3 70B Instruct (free)
3.9
Gemini 1.5 Pro (May 2024)
12.5
WeirdML
Gemini 1.5 Pro (May 2024) leads by +7.8
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Llama 3.3 70B Instruct (free)
14.4
Gemini 1.5 Pro (May 2024)
22.2
Full benchmark table
| Benchmark | Llama 3.3 70B Instruct (free) | Gemini 1.5 Pro (May 2024) |
|---|---|---|
Balrog Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning. | 23.0 | 21.0 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 29.9 | 27.8 |
MATH level 5 MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics. | 41.6 | 40.8 |
MMLU Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge. | 81.7 | 81.2 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 5.0 | 6.7 |
SimpleBench SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking. | 3.9 | 12.5 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 14.4 | 22.2 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.00 | $0.00 | 131K tokens (~66 books) | — | |
| — | — | — | — |