Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...
Tested on 15 benchmarks with 51.7% average. Top scores: Chatbot Arena Elo — Coding (1513.5%), Chatbot Arena Elo — Overall (1460.2%), OTIS Mock AIME 2024-2025 (96.1%).
gpt-oss-20b (free) scores 61.0 (99% as good) at $0.00/1M input · 100% cheaper
Real-world software engineering tasks from GitHub issues. Models must diagnose bugs and write patches that pass test suites. Human-verified subset of SWE-bench.
Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
Original research-level math problems created by professional mathematicians. Problems are unpublished and cannot be memorized.
Graduate-level science questions written by PhD experts. Diamond subset contains questions where experts disagree, testing deep understanding.
Simple factual questions with verified correct answers. Tests accuracy of basic knowledge retrieval. Low scores indicate hallucination.
Tactical chess puzzles testing pattern recognition and multi-move calculation. Measures strategic reasoning ability.
- Typemultimodal
- Context262K tokens (~131 books)
- ReleasedApr 2026
- LicenseOpen Source
- StatusActive
- Cost / Message~$0.006