Compare · ModelsLive · 2 picked · head to head

o1 vs Qwen3.6 Flash

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Qwen3.6 Flash wins 7 of 9 shared benchmarks. Leads in knowledge · general · math.

Category leads
knowledge·Qwen3.6 Flashgeneral·Qwen3.6 Flashmath·Qwen3.6 Flashreasoning·o1
Hype vs Reality
o1
#158 by perf·no signal
QUIET
Qwen3.6 Flash
#168 by perf·#2 by attention
OVERHYPED
Best value
55.6x better value than o1
o1
1.2 pts/$
$37.50/M
Qwen3.6 Flash
67.4 pts/$
$0.66/M
Vendor risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Alibaba Qwen logo
Alibaba (Qwen)
$293.0B·Tier 1
Low risk
Head to head
o1Qwen3.6 Flash
Chess Puzzles
Qwen3.6 Flash leads by +5.3
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
o1
10.6
Qwen3.6 Flash
15.8
Dtbench
Qwen3.6 Flash leads by +4.0
o1
57.8
Qwen3.6 Flash
61.8
FrontierMath-2025-02-28-Private
Qwen3.6 Flash leads by +1.0
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
o1
9.3
Qwen3.6 Flash
10.3
FrontierMath-Tiers-1-3-v2-Private
Qwen3.6 Flash leads by +7.7
o1
14.7
Qwen3.6 Flash
22.5
GPQA diamond
Qwen3.6 Flash leads by +8.8
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
o1
69.0
Qwen3.6 Flash
77.8
Lmca
Qwen3.6 Flash leads by +10.3
o1
26.2
Qwen3.6 Flash
36.5
OTIS Mock AIME 2024-2025
Qwen3.6 Flash leads by +11.1
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
o1
73.3
Qwen3.6 Flash
84.4
SimpleBench
o1 leads by +5.9
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
o1
28.1
Qwen3.6 Flash
22.2
SimpleQA Verified
o1 leads by +25.2
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
o1
41.1
Qwen3.6 Flash
15.9
Full benchmark table
Benchmarko1Qwen3.6 Flash
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
10.615.8
Dtbench
57.861.8
FrontierMath-2025-02-28-Private
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
9.310.3
FrontierMath-Tiers-1-3-v2-Private
14.722.5
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
69.077.8
Lmca
26.236.5
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
73.384.4
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
28.122.2
SimpleQA Verified
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
41.115.9
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
OpenAI logoo1$15.00$60.00200K tokens (~100 books)$262.50
Alibaba Qwen logoQwen3.6 Flash$0.19$1.131.0M tokens (~500 books)$4.22
People also compared