Compare · ModelsLive · 2 picked · head to head

Mixtral 8x7B vs Mixtral 8x7B Instruct

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Mixtral 8x7B wins 11 of 11 shared benchmarks. Leads in knowledge · math.

Category leads
knowledge·Mixtral 8x7Bmath·Mixtral 8x7B
Hype vs Reality
Mixtral 8x7B
#70 by perf·no signal
QUIET
Mixtral 8x7B Instruct
#71 by perf·no signal
QUIET
Best value
Mixtral 8x7B
no price
Mixtral 8x7B Instruct
107.0 pts/$
$0.54/M
Vendor risk
Mistral AI logo
Mistral AI
$14.0B·Tier 1
Medium risk
Mistral AI logo
Mistral AI
$14.0B·Tier 1
Medium risk
Head to head
Mixtral 8x7BMixtral 8x7B Instruct
ANLI
ANLI (Adversarial NLI) · adversarially constructed natural language inference dataset where each round targets weaknesses found in previous model generations.
Mixtral 8x7B
32.8
Mixtral 8x7B Instruct
32.8
ARC AI2
AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval.
Mixtral 8x7B
83.1
Mixtral 8x7B Instruct
83.1
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Mixtral 8x7B
7.5
Mixtral 8x7B Instruct
7.5
GSM8K
Grade School Math 8K · 8,500 linguistically diverse grade-school math word problems that require multi-step reasoning to solve.
Mixtral 8x7B
74.4
Mixtral 8x7B Instruct
74.4
HellaSwag
HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.
Mixtral 8x7B
82.3
Mixtral 8x7B Instruct
82.3
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Mixtral 8x7B
9.9
Mixtral 8x7B Instruct
9.9
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Mixtral 8x7B
60.8
Mixtral 8x7B Instruct
60.8
OpenBookQA
OpenBookQA · science questions that require combining a given core fact with broad common knowledge, mimicking an open-book exam setting.
Mixtral 8x7B
81.1
Mixtral 8x7B Instruct
81.1
PIQA
PIQA (Physical Interaction QA) · tests intuitive physical reasoning by asking models to select the correct approach for everyday physical tasks.
Mixtral 8x7B
67.2
Mixtral 8x7B Instruct
67.2
TriviaQA
TriviaQA · reading comprehension benchmark with trivia questions, requiring models to find and reason over evidence from provided documents.
Mixtral 8x7B
82.2
Mixtral 8x7B Instruct
82.2
Winogrande
WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs.
Mixtral 8x7B
54.4
Mixtral 8x7B Instruct
54.4
Full benchmark table
BenchmarkMixtral 8x7BMixtral 8x7B Instruct
ANLI
ANLI (Adversarial NLI) · adversarially constructed natural language inference dataset where each round targets weaknesses found in previous model generations.
32.832.8
ARC AI2
AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval.
83.183.1
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
7.57.5
GSM8K
Grade School Math 8K · 8,500 linguistically diverse grade-school math word problems that require multi-step reasoning to solve.
74.474.4
HellaSwag
HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.
82.382.3
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
9.99.9
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
60.860.8
OpenBookQA
OpenBookQA · science questions that require combining a given core fact with broad common knowledge, mimicking an open-book exam setting.
81.181.1
PIQA
PIQA (Physical Interaction QA) · tests intuitive physical reasoning by asking models to select the correct approach for everyday physical tasks.
67.267.2
TriviaQA
TriviaQA · reading comprehension benchmark with trivia questions, requiring models to find and reason over evidence from provided documents.
82.282.2
Winogrande
WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs.
54.454.4
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Mistral AI logoMixtral 8x7B
Mistral AI logoMixtral 8x7B Instruct$0.54$0.5433K tokens (~16 books)$5.40