Compare · ModelsLive · 2 picked · head to head

Claude 3 Opus vs GPT-5.4 Nano

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

GPT-5.4 Nano wins 6 of 7 shared benchmarks. Leads in knowledge · general · math.

Category leads
knowledge·GPT-5.4 Nanogeneral·GPT-5.4 Nanomath·GPT-5.4 Nanocoding·GPT-5.4 Nano
Hype vs Reality
Claude 3 Opus
#261 by perf·no signal
QUIET
GPT-5.4 Nano
#237 by perf·#4 by attention
OVERHYPED
Best value
Claude 3 Opus
n/a
no price
GPT-5.4 Nano
45.1 pts/$
$0.72/M
Vendor risk
Anthropic logo
Anthropic
$965.0B·Tier 1
Medium risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Head to head
Claude 3 OpusGPT-5.4 Nano
Chess Puzzles
GPT-5.4 Nano leads by +26.3
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Claude 3 Opus
0.0
GPT-5.4 Nano
26.4
Dtbench
GPT-5.4 Nano leads by +31.1
Claude 3 Opus
36.0
GPT-5.4 Nano
67.1
GPQA diamond
GPT-5.4 Nano leads by +41.7
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Claude 3 Opus
29.6
GPT-5.4 Nano
71.3
Lmca
GPT-5.4 Nano leads by +23.4
Claude 3 Opus
20.0
GPT-5.4 Nano
43.4
OTIS Mock AIME 2024-2025
GPT-5.4 Nano leads by +83.1
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Claude 3 Opus
4.6
GPT-5.4 Nano
87.8
SimpleQA Verified
Claude 3 Opus leads by +0.9
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Claude 3 Opus
12.6
GPT-5.4 Nano
11.7
WeirdML
GPT-5.4 Nano leads by +30.0
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Claude 3 Opus
19.2
GPT-5.4 Nano
49.2
Full benchmark table
BenchmarkClaude 3 OpusGPT-5.4 Nano
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
0.026.4
Dtbench
36.067.1
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
29.671.3
Lmca
20.043.4
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
4.687.8
SimpleQA Verified
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
12.611.7
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
19.249.2
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Anthropic logoClaude 3 Opus————
OpenAI logoGPT-5.4 Nano$0.20$1.25400K tokens (~200 books)$4.63