Compare · ModelsLive · 2 picked · head to head
R1 vs Qwen3.6 Flash
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Qwen3.6 Flash wins on 3/4 benchmarks
Qwen3.6 Flash wins 3 of 4 shared benchmarks. Leads in knowledge · math · reasoning.
Category leads
knowledge·Qwen3.6 Flashmath·Qwen3.6 Flashreasoning·Qwen3.6 Flash
Hype vs Reality
Attention vs performance
R1
#156 by perf·#14 by attention
Qwen3.6 Flash
#168 by perf·#2 by attention
Best value
Qwen3.6 Flash
2.4x better value than R1
R1
28.5 pts/$
$1.60/M
Qwen3.6 Flash
67.4 pts/$
$0.66/M
Vendor risk
Mixed exposure
One or more vendors flagged
DeepSeek
$3.4B·Tier 1
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
4 benchmarks · 2 models
R1Qwen3.6 Flash
GPQA diamond
Qwen3.6 Flash leads by +15.5
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
R1
62.3
Qwen3.6 Flash
77.8
OTIS Mock AIME 2024-2025
Qwen3.6 Flash leads by +31.1
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
R1
53.3
Qwen3.6 Flash
84.4
SimpleBench
Qwen3.6 Flash leads by +5.2
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
R1
17.1
Qwen3.6 Flash
22.2
SimpleQA Verified
R1 leads by +11.5
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
R1
27.4
Qwen3.6 Flash
15.9
Full benchmark table
| Benchmark | R1 | Qwen3.6 Flash |
|---|---|---|
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 62.3 | 77.8 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 53.3 | 84.4 |
SimpleBench SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking. | 17.1 | 22.2 |
SimpleQA Verified SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information. | 27.4 | 15.9 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.70 | $2.50 | 64K tokens (~32 books) | $11.50 | |
| $0.19 | $1.13 | 1.0M tokens (~500 books) | $4.22 |