Compare · ModelsLive · 2 picked · head to head
Qwen3.6 Plus vs GPT-5 Mini
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Qwen3.6 Plus wins on 11/14 benchmarks
Qwen3.6 Plus wins 11 of 14 shared benchmarks. Leads in math · knowledge · coding.
Category leads
math·Qwen3.6 Plusknowledge·Qwen3.6 Pluscoding·Qwen3.6 Plusreasoning·Qwen3.6 Pluslanguage·GPT-5 Mini
Hype vs Reality
Attention vs performance
Qwen3.6 Plus
#55 by perf·no signal
GPT-5 Mini
#90 by perf·no signal
Best value
Qwen3.6 Plus
1.1x better value than GPT-5 Mini
Qwen3.6 Plus
52.7 pts/$
$1.14/M
GPT-5 Mini
48.3 pts/$
$1.13/M
Vendor risk
Who is behind the model
Alibaba (Qwen)
$293.0B·Tier 1
OpenAI
$840.0B·Tier 1
Head to head
14 benchmarks · 2 models
Qwen3.6 PlusGPT-5 Mini
FrontierMath-2025-02-28-Private
GPT-5 Mini leads by +1.0
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Qwen3.6 Plus
26.2
GPT-5 Mini
27.2
FrontierMath-Tier-4-2025-07-01-Private
Qwen3.6 Plus leads by +2.1
FrontierMath Tier 4 (Jul 2025) · the most challenging tier of frontier mathematics, containing problems that push the absolute limits of AI mathematical reasoning.
Qwen3.6 Plus
8.3
GPT-5 Mini
6.3
GPQA diamond
Qwen3.6 Plus leads by +16.5
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Qwen3.6 Plus
83.2
GPT-5 Mini
66.7
LiveBench · Agentic Coding
Qwen3.6 Plus leads by +20.0
Qwen3.6 Plus
55.0
GPT-5 Mini
35.0
LiveBench · Coding
Qwen3.6 Plus leads by +2.1
Qwen3.6 Plus
78.2
GPT-5 Mini
76.1
LiveBench · Data Analysis
Qwen3.6 Plus leads by +20.3
Qwen3.6 Plus
69.9
GPT-5 Mini
49.6
LiveBench · If
GPT-5 Mini leads by +5.9
Qwen3.6 Plus
58.3
GPT-5 Mini
64.2
LiveBench · Language
Qwen3.6 Plus leads by +5.8
Qwen3.6 Plus
75.0
GPT-5 Mini
69.2
LiveBench · Mathematics
Qwen3.6 Plus leads by +9.3
Qwen3.6 Plus
83.7
GPT-5 Mini
74.4
LiveBench · Overall
Qwen3.6 Plus leads by +9.8
Qwen3.6 Plus
70.8
GPT-5 Mini
61.0
LiveBench · Reasoning
Qwen3.6 Plus leads by +17.2
Qwen3.6 Plus
75.8
GPT-5 Mini
58.6
OTIS Mock AIME 2024-2025
Qwen3.6 Plus leads by +3.9
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Qwen3.6 Plus
90.5
GPT-5 Mini
86.7
SimpleQA Verified
Qwen3.6 Plus leads by +28.1
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Qwen3.6 Plus
49.1
GPT-5 Mini
21.0
SWE-Bench verified
GPT-5 Mini leads by +6.8
SWE-bench Verified · 500 human-validated tasks from 12 real Python repositories (Django, Flask, scikit-learn, sympy, and others). Each task requires the model to produce a git patch that resolves a real GitHub issue and passes the test suite. The verified subset eliminates ambiguous tasks from the original SWE-bench. Claude Mythos Preview leads at 93.9%, crossing 90% for the first time in 2026. Opus 4.6 scores 80.8%. The benchmark remains the most-cited evaluation for code-generation capability.
Qwen3.6 Plus
57.9
GPT-5 Mini
64.7
Full benchmark table
| Benchmark | Qwen3.6 Plus | GPT-5 Mini |
|---|---|---|
FrontierMath-2025-02-28-Private FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning. | 26.2 | 27.2 |
FrontierMath-Tier-4-2025-07-01-Private FrontierMath Tier 4 (Jul 2025) · the most challenging tier of frontier mathematics, containing problems that push the absolute limits of AI mathematical reasoning. | 8.3 | 6.3 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 83.2 | 66.7 |
LiveBench · Agentic Coding | 55.0 | 35.0 |
LiveBench · Coding | 78.2 | 76.1 |
LiveBench · Data Analysis | 69.9 | 49.6 |
LiveBench · If | 58.3 | 64.2 |
LiveBench · Language | 75.0 | 69.2 |
LiveBench · Mathematics | 83.7 | 74.4 |
LiveBench · Overall | 70.8 | 61.0 |
LiveBench · Reasoning | 75.8 | 58.6 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 90.5 | 86.7 |
SimpleQA Verified SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information. | 49.1 | 21.0 |
SWE-Bench verified SWE-bench Verified · 500 human-validated tasks from 12 real Python repositories (Django, Flask, scikit-learn, sympy, and others). Each task requires the model to produce a git patch that resolves a real GitHub issue and passes the test suite. The verified subset eliminates ambiguous tasks from the original SWE-bench. Claude Mythos Preview leads at 93.9%, crossing 90% for the first time in 2026. Opus 4.6 scores 80.8%. The benchmark remains the most-cited evaluation for code-generation capability. | 57.9 | 64.7 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.33 | $1.95 | 1.0M tokens (~500 books) | $7.31 | |
| $0.25 | $2.00 | 400K tokens (~200 books) | $6.88 |