Compare · ModelsLive · 3 picked · head to head

Gemma 4 31B vs GLM 5.1 vs GPT-5.1-Codex-Max

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

GLM 5.1 wins 11 of 16 shared benchmarks. Leads in reasoning · language · math.

Category leads
coding·GPT-5.1-Codex-Maxreasoning·GLM 5.1language·GLM 5.1math·GLM 5.1knowledge·GLM 5.1speed·GLM 5.1arena·GLM 5.1
Hype vs Reality
Gemma 4 31B
#101 by perf·#8 by attention
DESERVED
GLM 5.1
#84 by perf·#3 by attention
DESERVED
GPT-5.1-Codex-Max
#14 by perf·#4 by attention
DESERVED
Best value
9.0x better value than GLM 5.1
Gemma 4 31B
245.6 pts/$
$0.22/M
GLM 5.1
27.1 pts/$
$2.00/M
GPT-5.1-Codex-Max
12.8 pts/$
$5.63/M
Vendor risk
Google DeepMind logo
Google DeepMind
$4.20T·Tier 1
Low risk
z-ai logo
z-ai
private · undisclosed
Unknown
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Head to head
Gemma 4 31BGLM 5.1GPT-5.1-Codex-Max
LiveBench · Agentic Coding
GPT-5.1-Codex-Max leads by +1.7
Gemma 4 31B
40.0
GLM 5.1
55.0
GPT-5.1-Codex-Max
56.7
LiveBench · Coding
GPT-5.1-Codex-Max leads by +6.0
Gemma 4 31B
60.3
GLM 5.1
75.4
GPT-5.1-Codex-Max
81.4
LiveBench · Data Analysis
GLM 5.1 leads by +4.5
Gemma 4 31B
58.8
GLM 5.1
63.2
GPT-5.1-Codex-Max
54.9
LiveBench · If
GLM 5.1 leads by +0.9
Gemma 4 31B
67.6
GLM 5.1
68.5
GPT-5.1-Codex-Max
67.1
LiveBench · Language
GPT-5.1-Codex-Max leads by +3.6
Gemma 4 31B
71.3
GLM 5.1
71.8
GPT-5.1-Codex-Max
75.4
LiveBench · Mathematics
GLM 5.1 leads by +1.2
Gemma 4 31B
73.9
GLM 5.1
84.9
GPT-5.1-Codex-Max
83.7
LiveBench · Overall
GPT-5.1-Codex-Max leads by +1.8
Gemma 4 31B
61.6
GLM 5.1
70.2
GPT-5.1-Codex-Max
72.0
LiveBench · Reasoning
GPT-5.1-Codex-Max leads by +12.0
Gemma 4 31B
59.4
GLM 5.1
72.5
GPT-5.1-Codex-Max
84.6
Artificial Analysis · Quality Index
GLM 5.1 leads by +25.5
Gemma 4 31B
14.7
GLM 5.1
40.2
Chatbot Arena Elo · Coding
GLM 5.1 leads by +143.8
Gemma 4 31B
1364.7
GLM 5.1
1508.5
Chatbot Arena Elo · Overall
GLM 5.1 leads by +11.7
Gemma 4 31B
1452.8
GLM 5.1
1464.6
Chess Puzzles
GLM 5.1 leads by +14.7
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Gemma 4 31B
0.0
GLM 5.1
14.8
GPQA diamond
GLM 5.1 leads by +18.8
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Gemma 4 31B
67.7
GLM 5.1
86.5
OTIS Mock AIME 2024-2025
GLM 5.1 leads by +20.0
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Gemma 4 31B
73.3
GLM 5.1
93.3
SimpleQA Verified
GLM 5.1 leads by +23.6
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Gemma 4 31B
10.4
GLM 5.1
34.0
WeirdML
GLM 5.1 leads by +4.8
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Gemma 4 31B
52.3
GLM 5.1
57.1
Full benchmark table
BenchmarkGemma 4 31BGLM 5.1GPT-5.1-Codex-Max
LiveBench · Agentic Coding
40.055.056.7
LiveBench · Coding
60.375.481.4
LiveBench · Data Analysis
58.863.254.9
LiveBench · If
67.668.567.1
LiveBench · Language
71.371.875.4
LiveBench · Mathematics
73.984.983.7
LiveBench · Overall
61.670.272.0
LiveBench · Reasoning
59.472.584.6
Artificial Analysis · Quality Index
14.740.2—
Chatbot Arena Elo · Coding
1364.71508.5—
Chatbot Arena Elo · Overall
1452.81464.6—
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
0.014.8—
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
67.786.5—
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
73.393.3—
SimpleQA Verified
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
10.434.0—
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
52.357.1—
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Google DeepMind logoGemma 4 31B$0.09$0.34262K tokens (~131 books)$1.53
z-ai logoGLM 5.1$0.97$3.04205K tokens (~102 books)$14.83
OpenAI logoGPT-5.1-Codex-Max$1.25$10.00400K tokens (~200 books)$34.38