Compare · ModelsLive · 2 picked · head to head

R1 vs Qwen3.6 Flash

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Qwen3.6 Flash wins 3 of 4 shared benchmarks. Leads in knowledge · math · reasoning.

Category leads
knowledge·Qwen3.6 Flashmath·Qwen3.6 Flashreasoning·Qwen3.6 Flash
Hype vs Reality
R1
#156 by perf·#14 by attention
UNDERRATED
Qwen3.6 Flash
#168 by perf·#2 by attention
OVERHYPED
Best value
2.4x better value than R1
R1
28.5 pts/$
$1.60/M
Qwen3.6 Flash
67.4 pts/$
$0.66/M
Vendor risk
One or more vendors flagged
DeepSeek logo
DeepSeek
$3.4B·Tier 1
Higher risk
Alibaba Qwen logo
Alibaba (Qwen)
$293.0B·Tier 1
Low risk
Head to head
R1Qwen3.6 Flash
GPQA diamond
Qwen3.6 Flash leads by +15.5
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
R1
62.3
Qwen3.6 Flash
77.8
OTIS Mock AIME 2024-2025
Qwen3.6 Flash leads by +31.1
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
R1
53.3
Qwen3.6 Flash
84.4
SimpleBench
Qwen3.6 Flash leads by +5.2
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
R1
17.1
Qwen3.6 Flash
22.2
SimpleQA Verified
R1 leads by +11.5
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
R1
27.4
Qwen3.6 Flash
15.9
Full benchmark table
BenchmarkR1Qwen3.6 Flash
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
62.377.8
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
53.384.4
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
17.122.2
SimpleQA Verified
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
27.415.9
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
DeepSeek logoR1$0.70$2.5064K tokens (~32 books)$11.50
Alibaba Qwen logoQwen3.6 Flash$0.19$1.131.0M tokens (~500 books)$4.22
People also compared