Compare · ModelsLive · 2 picked · head to head
DeepSeek V3.2 vs gpt-oss-120b
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
DeepSeek V3.2 wins on 18/25 benchmarks
DeepSeek V3.2 wins 18 of 25 shared benchmarks. Leads in speed · coding · agentic.
Category leads
speed·DeepSeek V3.2coding·DeepSeek V3.2agentic·DeepSeek V3.2arena·DeepSeek V3.2knowledge·DeepSeek V3.2general·DeepSeek V3.2reasoning·DeepSeek V3.2language·gpt-oss-120bmath·gpt-oss-120b
Hype vs Reality
Attention vs performance
DeepSeek V3.2
#142 by perf·no signal
gpt-oss-120b
#127 by perf·no signal
Best value
gpt-oss-120b
3.5x better value than DeepSeek V3.2
DeepSeek V3.2
136.3 pts/$
$0.35/M
gpt-oss-120b
475.4 pts/$
$0.10/M
Vendor risk
Mixed exposure
One or more vendors flagged
DeepSeek
$3.4B·Tier 1
OpenAI
$840.0B·Tier 1
Head to head
25 benchmarks · 2 models
DeepSeek V3.2gpt-oss-120b
Artificial Analysis · Quality Index
DeepSeek V3.2 leads by +30.1
DeepSeek V3.2
41.7
gpt-oss-120b
11.6
Aider polyglot
DeepSeek V3.2 leads by +32.4
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
DeepSeek V3.2
74.2
gpt-oss-120b
41.8
APEX-Agents
DeepSeek V3.2 leads by +2.6
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments.
DeepSeek V3.2
7.0
gpt-oss-120b
4.4
Chatbot Arena Elo · Overall
DeepSeek V3.2 leads by +73.2
DeepSeek V3.2
1424.8
gpt-oss-120b
1351.6
Chess Puzzles
gpt-oss-120b leads by +6.3
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
DeepSeek V3.2
9.5
gpt-oss-120b
15.8
Dtbench
DeepSeek V3.2 leads by +15.5
DeepSeek V3.2
76.0
gpt-oss-120b
60.5
GPQA diamond
DeepSeek V3.2 leads by +10.2
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
DeepSeek V3.2
77.9
gpt-oss-120b
67.7
LiveBench · Agentic Coding
DeepSeek V3.2 leads by +30.0
DeepSeek V3.2
46.7
gpt-oss-120b
16.7
LiveBench · Coding
DeepSeek V3.2 leads by +15.5
DeepSeek V3.2
75.7
gpt-oss-120b
60.2
LiveBench · Data Analysis
DeepSeek V3.2 leads by +6.2
DeepSeek V3.2
45.0
gpt-oss-120b
38.8
LiveBench · If
gpt-oss-120b leads by +27.2
DeepSeek V3.2
23.1
gpt-oss-120b
50.3
LiveBench · Language
DeepSeek V3.2 leads by +15.6
DeepSeek V3.2
64.2
gpt-oss-120b
48.6
LiveBench · Mathematics
gpt-oss-120b leads by +4.9
DeepSeek V3.2
64.0
gpt-oss-120b
68.9
LiveBench · Overall
DeepSeek V3.2 leads by +5.8
DeepSeek V3.2
51.8
gpt-oss-120b
46.1
LiveBench · Reasoning
DeepSeek V3.2 leads by +5.0
DeepSeek V3.2
44.3
gpt-oss-120b
39.2
Lmca
DeepSeek V3.2 leads by +7.9
DeepSeek V3.2
33.9
gpt-oss-120b
26.1
OpenCompass · AIME2025
gpt-oss-120b leads by +0.4
DeepSeek V3.2
93.0
gpt-oss-120b
93.4
OpenCompass · GPQA-Diamond
DeepSeek V3.2 leads by +5.7
DeepSeek V3.2
84.6
gpt-oss-120b
78.9
OpenCompass · HLE
DeepSeek V3.2 leads by +4.9
DeepSeek V3.2
23.2
gpt-oss-120b
18.3
OpenCompass · IFEval
gpt-oss-120b leads by +0.5
DeepSeek V3.2
89.7
gpt-oss-120b
90.2
OpenCompass · LiveCodeBenchV6
gpt-oss-120b leads by +3.0
DeepSeek V3.2
75.4
gpt-oss-120b
78.4
OpenCompass · MMLU-Pro
DeepSeek V3.2 leads by +6.1
DeepSeek V3.2
85.8
gpt-oss-120b
79.7
OTIS Mock AIME 2024-2025
gpt-oss-120b leads by +1.1
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
DeepSeek V3.2
87.8
gpt-oss-120b
88.9
SimpleQA Verified
DeepSeek V3.2 leads by +13.6
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
DeepSeek V3.2
27.5
gpt-oss-120b
13.9
Terminal Bench
DeepSeek V3.2 leads by +20.8
Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence.
DeepSeek V3.2
39.5
gpt-oss-120b
18.7
Full benchmark table
| Benchmark | DeepSeek V3.2 | gpt-oss-120b |
|---|---|---|
Artificial Analysis · Quality Index | 41.7 | 11.6 |
Aider polyglot Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework. | 74.2 | 41.8 |
APEX-Agents APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments. | 7.0 | 4.4 |
Chatbot Arena Elo · Overall | 1424.8 | 1351.6 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 9.5 | 15.8 |
Dtbench | 76.0 | 60.5 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 77.9 | 67.7 |
LiveBench · Agentic Coding | 46.7 | 16.7 |
LiveBench · Coding | 75.7 | 60.2 |
LiveBench · Data Analysis | 45.0 | 38.8 |
LiveBench · If | 23.1 | 50.3 |
LiveBench · Language | 64.2 | 48.6 |
LiveBench · Mathematics | 64.0 | 68.9 |
LiveBench · Overall | 51.8 | 46.1 |
LiveBench · Reasoning | 44.3 | 39.2 |
Lmca | 33.9 | 26.1 |
OpenCompass · AIME2025 | 93.0 | 93.4 |
OpenCompass · GPQA-Diamond | 84.6 | 78.9 |
OpenCompass · HLE | 23.2 | 18.3 |
OpenCompass · IFEval | 89.7 | 90.2 |
OpenCompass · LiveCodeBenchV6 | 75.4 | 78.4 |
OpenCompass · MMLU-Pro | 85.8 | 79.7 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 87.8 | 88.9 |
SimpleQA Verified SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information. | 27.5 | 13.9 |
Terminal Bench Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence. | 39.5 | 18.7 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.28 | $0.42 | 164K tokens (~82 books) | $3.15 | |
| $0.04 | $0.17 | 131K tokens (~66 books) | $0.70 |