Gecko Tests · ScorecardAs of 2026-10-05 · 57 models tested · 13 with 3+ tests ranked
Gecko Scorecard · Who Passes Our Tests?
All of BenchGecko's own tests on one page, in plain words. Not how smart a model is, but how it behaves.
In plain wordsEach column is one test. A is best, E is worst. The Gecko Score adds them up: the higher, the better the model behaves on our tests.
Top 5 on the Gecko Tests
#1
Gemini 3.8 Flash
77
Gecko Score · 3 tests
Best at: draws the world from memory
#2
MiniMax M3
67
Gecko Score · 3 tests
Best at: fair price in other languages
#3
Claude Opus 5.5
65
Gecko Score · 4 tests
Best at: knows recent news
#4
Claude Sonnet 5.5
65
Gecko Score · 4 tests
Best at: knows recent news
#5
GPT-6 Luna
55
Gecko Score · 5 tests
Best at: draws the world from memory
| # | Model | Gecko Score | Who Are You Knows who made it · 57 models | World Map Draws the world from memory · 5 models | Censorship Index Answers normal questions · 9 models | Knowledge Horizon Knows recent news · 7 models | Tokenizer Tax Fair price in other languages · 54 models |
|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash | 77 | BKnows who made it | A97.4% of the map right | not tested yet | not tested yet | B44% more tokens outside English |
| 2 | MiniMax M3 | 67 | BKnows who made it | C92.8% of the map right | not tested yet | not tested yet | A36% more tokens outside English |
| 3 | Claude Opus 5.5 | 65 | BKnows who made it | not tested yet | AAnswers almost everything | AKnows news up to May 2026 | E73% more tokens outside English |
| 4 | Claude Sonnet 5.5 | 65 | BKnows who made it | not tested yet | AAnswers almost everything | AKnows news up to May 2026 | E73% more tokens outside English |
| 5 | GPT-6 Luna | 55 | BKnows who made it | B96.2% of the map right | CAnswers almost everything | CKnows news up to Apr 2026 | C56% more tokens outside English |
| 6 | GPT-6 Sol | 52 | DKnows who made it | not tested yet | AAnswers almost everything | CKnows news up to Apr 2026 | C53% more tokens outside English |
| 7 | GPT-6 Luna Pro | 49 | BKnows who made it | not tested yet | CAnswers almost everything | CKnows news up to Apr 2026 | D57% more tokens outside English |
| 8 | GPT-6.1 Sol Pro | 39 | BKnows who made it | not tested yet | CAnswers almost everything | EKnows news up to Mar 2026 | D64% more tokens outside English |
| 9 | GLM 5.3 Flash | 39 | BKnows who made it | D89.7% of the map right | not tested yet | not tested yet | D70% more tokens outside English |
| 10 | GPT-6.1 Sol | 39 | BKnows who made it | not tested yet | CAnswers almost everything | EKnows news up to Mar 2026 | not tested yet |
| 11 | Claude Opus 5 | 31 | BKnows who made it | not tested yet | EAnswers almost everything | not tested yet | E73% more tokens outside English |
| 12 | Llama 4 Maverick | 20 | ESometimes says it is Google or OpenAI | E79.2% of the map right | not tested yet | not tested yet | B44% more tokens outside English |
| 13 | DeepSeek V4.1 Flash | 12 | ESometimes says it is OpenAI | not tested yet | ERefuses or dodges 6% | not tested yet | D68% more tokens outside English |
| Claude Fable 5.1 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | E73% more tokens outside English | |
| Claude Opus 5 (Fast) | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | E73% more tokens outside English | |
| DeepSeek V4 Flash | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | A7% more tokens outside English | |
| DeepSeek V4 Flash 0731 | 2/3 tests | ESometimes says it is Google or Anthropic | not tested yet | not tested yet | not tested yet | D68% more tokens outside English | |
| DeepSeek V4 Pro | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | D68% more tokens outside English | |
| DeepSeek V4 Pro 0813 | 2/3 tests | ESometimes says it is OpenAI or Anthropic | not tested yet | not tested yet | not tested yet | A17% more tokens outside English | |
| Devstral 2 2512 | 2/3 tests | ESometimes says it is OpenAI | not tested yet | not tested yet | not tested yet | C53% more tokens outside English | |
| Gemini 3.5 Flash Lite | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | B44% more tokens outside English | |
| Gemini 3.6 Flash | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | B44% more tokens outside English | |
| Gemini 3.7 Flash | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | B44% more tokens outside English | |
| Gemma 4 26B A4B | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | B44% more tokens outside English | |
| GLM 5.1 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | D71% more tokens outside English | |
| GLM 5.3 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | E90% more tokens outside English | |
| GLM 5.3 FlashX | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | D70% more tokens outside English | |
| GLM 5.3 Prime | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | D70% more tokens outside English | |
| Grok 4.20 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | B44% more tokens outside English | |
| Grok 4.3 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | B44% more tokens outside English | |
| Grok 4.5 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | C50% more tokens outside English | |
| Grok 4.6 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | C50% more tokens outside English | |
| Grok 4.7 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | C50% more tokens outside English | |
| Grok Build 0.1 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | B44% more tokens outside English | |
| Kimi K2 0905 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | E92% more tokens outside English | |
| Kimi K2 Thinking | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | E92% more tokens outside English | |
| Kimi K2.5 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | E100% more tokens outside English | |
| Kimi K2.6 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | E92% more tokens outside English | |
| Kimi K2.7 Code | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | E92% more tokens outside English | |
| Kimi K3 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | E93% more tokens outside English | |
| Llama 4 Scout | 2/3 tests | DSometimes says it is OpenAI | not tested yet | not tested yet | not tested yet | B43% more tokens outside English | |
| MiniMax M2 | 2/3 tests | ESometimes says it is Anthropic | not tested yet | not tested yet | not tested yet | A36% more tokens outside English | |
| MiniMax M2-her | 2/3 tests | DKnows who made it | not tested yet | not tested yet | not tested yet | A36% more tokens outside English | |
| MiniMax M2.1 | 2/3 tests | ESometimes says it is Anthropic | not tested yet | not tested yet | not tested yet | A36% more tokens outside English | |
| MiniMax M2.5 | 2/3 tests | ESometimes says it is OpenAI | not tested yet | not tested yet | not tested yet | D64% more tokens outside English | |
| MiniMax M2.7 | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | A35% more tokens outside English | |
| Ministral 3 14B 2512 | 2/3 tests | DSometimes says it is OpenAI | not tested yet | not tested yet | not tested yet | C53% more tokens outside English | |
| Ministral 3 3B 2512 | 2/3 tests | ESometimes says it is OpenAI | not tested yet | not tested yet | not tested yet | C53% more tokens outside English | |
| Mistral Medium 3.5 | 2/3 tests | ESometimes says it is OpenAI | not tested yet | not tested yet | not tested yet | C53% more tokens outside English | |
| Mistral Small 4 | 2/3 tests | DSometimes says it is OpenAI | not tested yet | not tested yet | not tested yet | C53% more tokens outside English | |
| Qwen3.8 2.4T A95B | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | A39% more tokens outside English | |
| Qwen3.8 27B | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | B41% more tokens outside English | |
| Qwen3.8 Flash | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | A35% more tokens outside English | |
| Qwen3.8 Max (0902) | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | A35% more tokens outside English | |
| Qwen3.8 Max Prime | 2/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | A35% more tokens outside English | |
| Gemini 3.5 Flash | 1/3 tests | BKnows who made it | not tested yet | not tested yet | not tested yet | not tested yet | |
| GLM 5.2 | 1/3 tests | DSometimes says it is Google | not tested yet | not tested yet | not tested yet | not tested yet |
Every new model from a tracked lab is tested automatically within a day of release; grades update as models are added. Method per test on each test page.
How to cite · data as of 2026-10-05
BenchGecko Gecko Scorecard. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/gecko-tests/scorecard
BenchGecko measures this data itself: free to reuse under CC BY 4.0 with the credit "Source: BenchGecko" and a link. JSON · llms.txt · MCP
Frequently Asked Questions
The average of a model's rank across BenchGecko's own tests (0 to 100, higher is better). Each test is ranked separately, so a model is compared with the other models that took the same test. A model needs at least 3 tests to get a score.