Compare · ModelsLive · 2 picked · head to head
Qwen3 235B A22B vs Kimi K2.5
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Kimi K2.5 wins on 4/4 benchmarks
Kimi K2.5 wins 4 of 4 shared benchmarks. Leads in knowledge · reasoning · coding.
Category leads
knowledge·Kimi K2.5reasoning·Kimi K2.5coding·Kimi K2.5
Hype vs Reality
Attention vs performance
Qwen3 235B A22B
#76 by perf·no signal
Kimi K2.5
#104 by perf·no signal
Best value
Qwen3 235B A22B
1.1x better value than Kimi K2.5
Qwen3 235B A22B
49.6 pts/$
$1.14/M
Kimi K2.5
43.3 pts/$
$1.20/M
Vendor risk
Who is behind the model
Alibaba (Qwen)
$293.0B·Tier 1
moonshotai
private · undisclosed
Head to head
4 benchmarks · 2 models
Qwen3 235B A22BKimi K2.5
Fiction.LiveBench
Kimi K2.5 leads by +18.4
Fiction.LiveBench · a continuously updated benchmark using recently published fiction to test reading comprehension and reasoning, preventing data contamination.
Qwen3 235B A22B
67.7
Kimi K2.5
86.1
GPQA diamond
Kimi K2.5 leads by +22.5
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Qwen3 235B A22B
60.9
Kimi K2.5
83.5
SimpleBench
Kimi K2.5 leads by +19.0
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Qwen3 235B A22B
17.2
Kimi K2.5
36.2
WeirdML
Kimi K2.5 leads by +8.3
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Qwen3 235B A22B
37.3
Kimi K2.5
45.6
Full benchmark table
| Benchmark | Qwen3 235B A22B | Kimi K2.5 |
|---|---|---|
Fiction.LiveBench Fiction.LiveBench · a continuously updated benchmark using recently published fiction to test reading comprehension and reasoning, preventing data contamination. | 67.7 | 86.1 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 60.9 | 83.5 |
SimpleBench SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking. | 17.2 | 36.2 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 37.3 | 45.6 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.46 | $1.82 | 131K tokens (~66 books) | $7.96 | |
| $0.38 | $2.02 | 262K tokens (~131 books) | $7.88 |