Compare · ModelsLive · 2 picked · head to head
Nemotron 3 Ultra vs Qwen3.7 Plus
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Qwen3.7 Plus wins on 12/19 benchmarks
Qwen3.7 Plus wins 12 of 19 shared benchmarks. Leads in speed · knowledge · math.
Category leads
speed·Qwen3.7 Plusknowledge·Qwen3.7 Plusgeneral·Nemotron 3 Ultracoding·Nemotron 3 Ultramath·Qwen3.7 Plus
Hype vs Reality
Attention vs performance
Nemotron 3 Ultra
#148 by perf·#16 by attention
Qwen3.7 Plus
#183 by perf·#2 by attention
Best value
Qwen3.7 Plus
1.5x better value than Nemotron 3 Ultra
Nemotron 3 Ultra
34.3 pts/$
$1.35/M
Qwen3.7 Plus
51.5 pts/$
$0.80/M
Vendor risk
Who is behind the model
NVIDIA
$5.64T·Big Tech
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
19 benchmarks · 2 models
Nemotron 3 UltraQwen3.7 Plus
Artificial Analysis · Agentic Index
Nemotron 3 Ultra leads by +6.6
Artificial Analysis Agentic Index · a composite score measuring how well a model performs in agentic workflows · multi-step tool use, planning, error recovery, and autonomous task completion. Aggregates results from multiple agentic benchmarks including SWE-bench, tool-use tests, and planning evaluations. The canonical single-number metric for "how good is this model as an agent?"
Nemotron 3 Ultra
27.4
Qwen3.7 Plus
20.8
Artificial Analysis · Coding Index
Qwen3.7 Plus leads by +6.6
Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads.
Nemotron 3 Ultra
49.3
Qwen3.7 Plus
55.9
Artificial Analysis · CritPt
Qwen3.7 Plus leads by +6.0
Nemotron 3 Ultra
3.1
Qwen3.7 Plus
9.1
Artificial Analysis · GDPval
Nemotron 3 Ultra leads by +12.2
Nemotron 3 Ultra
25.0
Qwen3.7 Plus
12.8
Artificial Analysis · GPQA Diamond
Qwen3.7 Plus leads by +3.3
Nemotron 3 Ultra
86.7
Qwen3.7 Plus
90.0
Artificial Analysis · Humanity's Last Exam
Qwen3.7 Plus leads by +7.2
Nemotron 3 Ultra
28.4
Qwen3.7 Plus
35.6
Artificial Analysis · IFBench
Nemotron 3 Ultra leads by +3.4
Nemotron 3 Ultra
81.4
Qwen3.7 Plus
78.0
Artificial Analysis · Long Context Reasoning
Nemotron 3 Ultra leads by +6.3
Nemotron 3 Ultra
79.3
Qwen3.7 Plus
73.0
Artificial Analysis · Quality Index
Qwen3.7 Plus leads by +2.2
Nemotron 3 Ultra
22.9
Qwen3.7 Plus
25.2
Artificial Analysis · SciCode
Qwen3.7 Plus leads by +5.8
Nemotron 3 Ultra
40.3
Qwen3.7 Plus
46.1
Artificial Analysis · tau2-Bench Telecom
Qwen3.7 Plus leads by +9.7
Nemotron 3 Ultra
83.3
Qwen3.7 Plus
93.0
Artificial Analysis · Terminal-Bench Hard
Qwen3.7 Plus leads by +10.6
Nemotron 3 Ultra
36.4
Qwen3.7 Plus
47.0
Chess Puzzles
Qwen3.7 Plus leads by +12.6
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Nemotron 3 Ultra
7.4
Qwen3.7 Plus
20.0
Dtbench
Nemotron 3 Ultra leads by +10.2
Nemotron 3 Ultra
83.5
Qwen3.7 Plus
73.3
Frontiercode
Nemotron 3 Ultra leads by +3.4
Nemotron 3 Ultra
13.6
Qwen3.7 Plus
10.2
GPQA diamond
Qwen3.7 Plus leads by +3.4
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Nemotron 3 Ultra
80.5
Qwen3.7 Plus
83.8
Lmca
Qwen3.7 Plus leads by +0.8
Nemotron 3 Ultra
43.4
Qwen3.7 Plus
44.2
Mystery Game Puzzles
Nemotron 3 Ultra leads by +3.3
Nemotron 3 Ultra
11.9
Qwen3.7 Plus
8.6
OTIS Mock AIME 2024-2025
Qwen3.7 Plus leads by +6.7
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Nemotron 3 Ultra
86.7
Qwen3.7 Plus
93.3
Full benchmark table
| Benchmark | Nemotron 3 Ultra | Qwen3.7 Plus |
|---|---|---|
Artificial Analysis · Agentic Index Artificial Analysis Agentic Index · a composite score measuring how well a model performs in agentic workflows · multi-step tool use, planning, error recovery, and autonomous task completion. Aggregates results from multiple agentic benchmarks including SWE-bench, tool-use tests, and planning evaluations. The canonical single-number metric for "how good is this model as an agent?" | 27.4 | 20.8 |
Artificial Analysis · Coding Index Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads. | 49.3 | 55.9 |
Artificial Analysis · CritPt | 3.1 | 9.1 |
Artificial Analysis · GDPval | 25.0 | 12.8 |
Artificial Analysis · GPQA Diamond | 86.7 | 90.0 |
Artificial Analysis · Humanity's Last Exam | 28.4 | 35.6 |
Artificial Analysis · IFBench | 81.4 | 78.0 |
Artificial Analysis · Long Context Reasoning | 79.3 | 73.0 |
Artificial Analysis · Quality Index | 22.9 | 25.2 |
Artificial Analysis · SciCode | 40.3 | 46.1 |
Artificial Analysis · tau2-Bench Telecom | 83.3 | 93.0 |
Artificial Analysis · Terminal-Bench Hard | 36.4 | 47.0 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 7.4 | 20.0 |
Dtbench | 83.5 | 73.3 |
Frontiercode | 13.6 | 10.2 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 80.5 | 83.8 |
Lmca | 43.4 | 44.2 |
Mystery Game Puzzles | 11.9 | 8.6 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 86.7 | 93.3 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.50 | $2.20 | 1.0M tokens (~500 books) | $9.25 | |
| $0.32 | $1.28 | 1.0M tokens (~500 books) | $5.60 |