Compare · ModelsLive · 2 picked · head to head

DeepSeek V3.2 vs gpt-oss-120b

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

DeepSeek V3.2 wins 12 of 16 shared benchmarks. Leads in coding · agentic · arena.

Category leads
coding·DeepSeek V3.2agentic·DeepSeek V3.2arena·DeepSeek V3.2knowledge·DeepSeek V3.2reasoning·DeepSeek V3.2language·gpt-oss-120bmath·gpt-oss-120b
Hype vs Reality
DeepSeek V3.2
#108 by perf·no signal
QUIET
gpt-oss-120b
#137 by perf·no signal
QUIET
Best value
2.3x better value than DeepSeek V3.2
DeepSeek V3.2
191.0 pts/$
$0.27/M
gpt-oss-120b
434.3 pts/$
$0.11/M
Vendor risk
One or more vendors flagged
DeepSeek logo
DeepSeek
$3.4B·Tier 1
Higher risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Head to head
DeepSeek V3.2gpt-oss-120b
Aider polyglot
DeepSeek V3.2 leads by +32.4
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
DeepSeek V3.2
74.2
gpt-oss-120b
41.8
APEX-Agents
DeepSeek V3.2 leads by +2.3
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments.
DeepSeek V3.2
7.0
gpt-oss-120b
4.7
Chatbot Arena Elo · Overall
DeepSeek V3.2 leads by +72.3
DeepSeek V3.2
1425.0
gpt-oss-120b
1352.7
Chess Puzzles
gpt-oss-120b leads by +6.0
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
DeepSeek V3.2
14.0
gpt-oss-120b
20.0
GPQA diamond
DeepSeek V3.2 leads by +10.2
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
DeepSeek V3.2
77.9
gpt-oss-120b
67.7
LiveBench · Agentic Coding
DeepSeek V3.2 leads by +30.0
DeepSeek V3.2
46.7
gpt-oss-120b
16.7
LiveBench · Coding
DeepSeek V3.2 leads by +15.5
DeepSeek V3.2
75.7
gpt-oss-120b
60.2
LiveBench · Data Analysis
DeepSeek V3.2 leads by +6.2
DeepSeek V3.2
45.0
gpt-oss-120b
38.8
LiveBench · If
gpt-oss-120b leads by +27.2
DeepSeek V3.2
23.1
gpt-oss-120b
50.3
LiveBench · Language
DeepSeek V3.2 leads by +15.6
DeepSeek V3.2
64.2
gpt-oss-120b
48.6
LiveBench · Mathematics
gpt-oss-120b leads by +4.9
DeepSeek V3.2
64.0
gpt-oss-120b
68.9
LiveBench · Overall
DeepSeek V3.2 leads by +5.8
DeepSeek V3.2
51.8
gpt-oss-120b
46.1
LiveBench · Reasoning
DeepSeek V3.2 leads by +5.0
DeepSeek V3.2
44.3
gpt-oss-120b
39.2
OTIS Mock AIME 2024-2025
gpt-oss-120b leads by +1.1
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
DeepSeek V3.2
87.8
gpt-oss-120b
88.9
SimpleQA Verified
DeepSeek V3.2 leads by +13.6
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
DeepSeek V3.2
27.5
gpt-oss-120b
13.9
Terminal Bench
DeepSeek V3.2 leads by +20.9
Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence.
DeepSeek V3.2
39.6
gpt-oss-120b
18.7
Full benchmark table
BenchmarkDeepSeek V3.2gpt-oss-120b
Aider polyglot
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
74.241.8
APEX-Agents
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments.
7.04.7
Chatbot Arena Elo · Overall
1425.01352.7
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
14.020.0
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
77.967.7
LiveBench · Agentic Coding
46.716.7
LiveBench · Coding
75.760.2
LiveBench · Data Analysis
45.038.8
LiveBench · If
23.150.3
LiveBench · Language
64.248.6
LiveBench · Mathematics
64.068.9
LiveBench · Overall
51.846.1
LiveBench · Reasoning
44.339.2
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
87.888.9
SimpleQA Verified
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
27.513.9
Terminal Bench
Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence.
39.618.7
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
DeepSeek logoDeepSeek V3.2$0.21$0.32131K tokens (~66 books)$2.41
OpenAI logogpt-oss-120b$0.04$0.18131K tokens (~66 books)$0.72