Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
Tested on 20 benchmarks with 59.9% average. Top scores: Chatbot Arena Elo — Coding (1461.9%), Chatbot Arena Elo — Overall (1444.2%), OTIS Mock AIME 2024-2025 (90.5%).
Regularly refreshed coding problems that avoid data contamination. New problems added monthly to prevent memorization.
Real-world software engineering tasks from GitHub issues. Models must diagnose bugs and write patches that pass test suites. Human-verified subset of SWE-bench.
LiveBench coding tasks that require multi-step reasoning and tool use. Tests planning and execution of complex coding workflows.
Regularly refreshed reasoning problems testing logical deduction, spatial reasoning, and analytical thinking.
Fresh data analysis tasks testing ability to interpret tables, charts, and statistical data.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
Regularly updated math problems that test numerical reasoning, algebra, calculus, and combinatorics.
Original research-level math problems created by professional mathematicians. Problems are unpublished and cannot be memorized.
- Typemultimodal
- Context1.0M tokens (~500 books)
- ReleasedApr 2026
- LicenseOpen Source
- StatusActive
- Cost / Message~$0.003