Gecko Tests · ScorecardAs of 2026-10-05 · 57 models tested · 13 with 3+ tests ranked

Gecko Scorecard · Who Passes Our Tests?

All of BenchGecko's own tests on one page, in plain words. Not how smart a model is, but how it behaves.

In plain wordsEach column is one test. A is best, E is worst. The Gecko Score adds them up: the higher, the better the model behaves on our tests.

Top 5 on the Gecko Tests

#ModelGecko ScoreWho Are You
Knows who made it · 57 models
World Map
Draws the world from memory · 5 models
Censorship Index
Answers normal questions · 9 models
Knowledge Horizon
Knows recent news · 7 models
Tokenizer Tax
Fair price in other languages · 54 models
1Gemini 3.8 Flash77
BKnows who made it
A97.4% of the map right
not tested yetnot tested yet
B44% more tokens outside English
2MiniMax M367
BKnows who made it
C92.8% of the map right
not tested yetnot tested yet
A36% more tokens outside English
3Claude Opus 5.565
BKnows who made it
not tested yet
AAnswers almost everything
AKnows news up to May 2026
E73% more tokens outside English
4Claude Sonnet 5.565
BKnows who made it
not tested yet
AAnswers almost everything
AKnows news up to May 2026
E73% more tokens outside English
5GPT-6 Luna55
BKnows who made it
B96.2% of the map right
CAnswers almost everything
CKnows news up to Apr 2026
C56% more tokens outside English
6GPT-6 Sol52
DKnows who made it
not tested yet
AAnswers almost everything
CKnows news up to Apr 2026
C53% more tokens outside English
7GPT-6 Luna Pro49
BKnows who made it
not tested yet
CAnswers almost everything
CKnows news up to Apr 2026
D57% more tokens outside English
8GPT-6.1 Sol Pro39
BKnows who made it
not tested yet
CAnswers almost everything
EKnows news up to Mar 2026
D64% more tokens outside English
9GLM 5.3 Flash39
BKnows who made it
D89.7% of the map right
not tested yetnot tested yet
D70% more tokens outside English
10GPT-6.1 Sol39
BKnows who made it
not tested yet
CAnswers almost everything
EKnows news up to Mar 2026
not tested yet
11Claude Opus 531
BKnows who made it
not tested yet
EAnswers almost everything
not tested yet
E73% more tokens outside English
12Llama 4 Maverick20
ESometimes says it is Google or OpenAI
E79.2% of the map right
not tested yetnot tested yet
B44% more tokens outside English
13DeepSeek V4.1 Flash12
ESometimes says it is OpenAI
not tested yet
ERefuses or dodges 6%
not tested yet
D68% more tokens outside English
Claude Fable 5.12/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
E73% more tokens outside English
Claude Opus 5 (Fast)2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
E73% more tokens outside English
DeepSeek V4 Flash2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
A7% more tokens outside English
DeepSeek V4 Flash 07312/3 tests
ESometimes says it is Google or Anthropic
not tested yetnot tested yetnot tested yet
D68% more tokens outside English
DeepSeek V4 Pro2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
D68% more tokens outside English
DeepSeek V4 Pro 08132/3 tests
ESometimes says it is OpenAI or Anthropic
not tested yetnot tested yetnot tested yet
A17% more tokens outside English
Devstral 2 25122/3 tests
ESometimes says it is OpenAI
not tested yetnot tested yetnot tested yet
C53% more tokens outside English
Gemini 3.5 Flash Lite2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
B44% more tokens outside English
Gemini 3.6 Flash2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
B44% more tokens outside English
Gemini 3.7 Flash2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
B44% more tokens outside English
Gemma 4 26B A4B 2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
B44% more tokens outside English
GLM 5.12/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
D71% more tokens outside English
GLM 5.32/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
E90% more tokens outside English
GLM 5.3 FlashX2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
D70% more tokens outside English
GLM 5.3 Prime2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
D70% more tokens outside English
Grok 4.202/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
B44% more tokens outside English
Grok 4.32/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
B44% more tokens outside English
Grok 4.52/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
C50% more tokens outside English
Grok 4.62/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
C50% more tokens outside English
Grok 4.72/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
C50% more tokens outside English
Grok Build 0.12/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
B44% more tokens outside English
Kimi K2 09052/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
E92% more tokens outside English
Kimi K2 Thinking2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
E92% more tokens outside English
Kimi K2.52/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
E100% more tokens outside English
Kimi K2.62/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
E92% more tokens outside English
Kimi K2.7 Code2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
E92% more tokens outside English
Kimi K32/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
E93% more tokens outside English
Llama 4 Scout2/3 tests
DSometimes says it is OpenAI
not tested yetnot tested yetnot tested yet
B43% more tokens outside English
MiniMax M22/3 tests
ESometimes says it is Anthropic
not tested yetnot tested yetnot tested yet
A36% more tokens outside English
MiniMax M2-her2/3 tests
DKnows who made it
not tested yetnot tested yetnot tested yet
A36% more tokens outside English
MiniMax M2.12/3 tests
ESometimes says it is Anthropic
not tested yetnot tested yetnot tested yet
A36% more tokens outside English
MiniMax M2.52/3 tests
ESometimes says it is OpenAI
not tested yetnot tested yetnot tested yet
D64% more tokens outside English
MiniMax M2.72/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
A35% more tokens outside English
Ministral 3 14B 25122/3 tests
DSometimes says it is OpenAI
not tested yetnot tested yetnot tested yet
C53% more tokens outside English
Ministral 3 3B 25122/3 tests
ESometimes says it is OpenAI
not tested yetnot tested yetnot tested yet
C53% more tokens outside English
Mistral Medium 3.52/3 tests
ESometimes says it is OpenAI
not tested yetnot tested yetnot tested yet
C53% more tokens outside English
Mistral Small 42/3 tests
DSometimes says it is OpenAI
not tested yetnot tested yetnot tested yet
C53% more tokens outside English
Qwen3.8 2.4T A95B2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
A39% more tokens outside English
Qwen3.8 27B2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
B41% more tokens outside English
Qwen3.8 Flash2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
A35% more tokens outside English
Qwen3.8 Max (0902)2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
A35% more tokens outside English
Qwen3.8 Max Prime2/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yet
A35% more tokens outside English
Gemini 3.5 Flash1/3 tests
BKnows who made it
not tested yetnot tested yetnot tested yetnot tested yet
GLM 5.21/3 tests
DSometimes says it is Google
not tested yetnot tested yetnot tested yetnot tested yet

Every new model from a tracked lab is tested automatically within a day of release; grades update as models are added. Method per test on each test page.

How to cite · data as of 2026-10-05

BenchGecko Gecko Scorecard. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/gecko-tests/scorecard

BenchGecko measures this data itself: free to reuse under CC BY 4.0 with the credit "Source: BenchGecko" and a link. JSON · llms.txt · MCP

The average of a model's rank across BenchGecko's own tests (0 to 100, higher is better). Each test is ranked separately, so a model is compared with the other models that took the same test. A model needs at least 3 tests to get a score.