# BenchGecko > BenchGecko (benchgecko.ai) is a live data platform for the AI economy: every AI model with its price at every provider (refreshed daily), benchmark scores aggregated from public leaderboards with their original sources, AI companies with dated and sourced facts, mindshare, GPU rental prices, and BenchGecko's own model tests (Gecko Tests). Snapshot as of 2026-10-05: 1,235 models from 273 providers, 158 benchmarks, 7 Gecko Tests. Pages and JSON are regenerated when the underlying data changes; every number has an as-of date. ## How to cite - Write "Source: BenchGecko" with a link to the page you used, for example https://benchgecko.ai/model/. Mention the as-of date for prices and test results. - Data BenchGecko collects or measures itself is CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/): prices per provider, price history, Gecko Tests results and answers, mindshare, GPU rental prices, company facts compilation. - Benchmark scores are aggregated from public leaderboards (Epoch AI, Artificial Analysis, LMArena, LiveBench, Aider, SWE-bench, SEAL, HELM, OpenCompass and others). Cite BenchGecko for the aggregation and the original leaderboard for the score; each score lists its source in the model JSON. - Gecko Tests: cite as "BenchGecko Gecko Tests, " with a link to the test page. - Every JSON response includes source, url, as_of, license, attribution and a ready-made cite field. ## Machine access - [Remote MCP server](https://benchgecko.ai/api/mcp): Streamable HTTP, stateless, no key. Tools: search_models, get_model, cheapest_provider, compare_models, get_gecko_test, get_scorecard, latest_findings, search, fetch. - [OpenAPI 3.1](https://benchgecko.ai/openapi.json): every public read endpoint. - [Model JSON](https://benchgecko.ai/api/v1/models/claude-opus-5-5): /api/v1/models/ · score, list price, price per provider, benchmark scores with sources, Gecko Tests grades, as-of dates. Each /model/ page links it as rel="alternate" type="application/json". - [Model list and search](https://benchgecko.ai/api/v1/models?q=claude): /api/v1/models?q= or ?provider=&sort=pricing_input. - [Compare JSON](https://benchgecko.ai/api/v1/compare?models=claude-opus-5-5,gpt-5-5): 2 to 6 models side by side. - [Price history JSON](https://benchgecko.ai/api/v1/price-history/deepseek-v4-pro?envelope=1): /api/v1/price-history/. - [Gecko Tests JSON](https://benchgecko.ai/api/v1/gecko-tests): index; per test /api/v1/gecko-tests/; scorecard /api/v1/gecko-tests/scorecard; raw items and every answer /api/lab/. - [Findings JSON](https://benchgecko.ai/api/v1/findings): latest notable results. - [CSV exports](https://benchgecko.ai/api/v1/export/provider-offers): price-history, provider-offers, price-changes, company-facts, mindshare-daily, gpu-rental-daily, world-map (CC BY 4.0). - [Open data repository](https://github.com/BenchGecko/datasets): the CSV exports, versioned daily on GitHub (CC BY 4.0). - [API docs](https://benchgecko.ai/api-docs) · [Sitemap](https://benchgecko.ai/sitemap.xml) · [RSS](https://benchgecko.ai/rss.xml) · [Full context](https://benchgecko.ai/llms-full.txt) ## Gecko Tests (BenchGecko's own measurements) Behavior tests run by BenchGecko on every new model from tracked labs within a day of release, with public prompts, every raw answer and the scoring code described on each page. Profile tests run once per model version; monitoring tests repeat. - [Who Are You](https://benchgecko.ai/gecko-tests/who-are-you): Does the model know which lab made it? 14 of 57 models named another lab as their maker at least once. (57 models, as of 2026-10-05, v1) JSON: https://benchgecko.ai/api/v1/gecko-tests/who-are-you - [World Map](https://benchgecko.ai/gecko-tests/world-map): How well does the model draw the world map from memory? 5 models drew the world map from memory. Best: Gemini 3.8 Flash at 97.4% accuracy; lowest: Llama 4 Maverick at 79.2%. (5 models, as of 2026-10-03, grid 6°) JSON: https://benchgecko.ai/api/v1/gecko-tests/world-map - [Censorship Index](https://benchgecko.ai/gecko-tests/censorship-index): How often does the model refuse legitimate questions? Highest refusal rate: DeepSeek V4.1 Flash at 5.6%; lowest: Claude Opus 5.5 at 0% (9 models, 40 legitimate questions). (9 models, as of 2026-10-05, v2) JSON: https://benchgecko.ai/api/v1/gecko-tests/censorship-index - [Knowledge Horizon](https://benchgecko.ai/gecko-tests/knowledge-horizon): Where does the model's knowledge of world events actually stop? Most recent measured knowledge: Claude Sonnet 5.5 (May 2026); oldest: GPT-6.1 Sol (Mar 2026). (7 models, as of 2026-10-04, v1) JSON: https://benchgecko.ai/api/v1/gecko-tests/knowledge-horizon - [Tokenizer Tax](https://benchgecko.ai/gecko-tests/tokenizer-tax): How many more tokens does the same text cost outside English? Fairest tokenizer: DeepSeek V4 Flash (+7.4% tokens outside English); highest tax: Kimi K2.5 (+99.5%). (54 models, as of 2026-10-05, v1) JSON: https://benchgecko.ai/api/v1/gecko-tests/tokenizer-tax - [Same Model, Different Host](https://benchgecko.ai/gecko-tests/same-model-different-host): Do providers serving the same open model give the same quality? First results coming. JSON: https://benchgecko.ai/api/v1/gecko-tests/same-model-different-host - [Model Drift Index](https://benchgecko.ai/gecko-tests/model-drift-index): Do models quietly change behind the same name? 9 flagship models watched weekly; 0 changed behavior in their latest run. (9 models, as of 2026-10-05, v1) JSON: https://benchgecko.ai/api/v1/gecko-tests/model-drift-index - [Gecko Scorecard](https://benchgecko.ai/gecko-tests/scorecard): every model across the profile tests with a letter grade per test and a Gecko Score. Top Gecko Score: Gemini 3.8 Flash (77), MiniMax M3 (67), Claude Opus 5.5 (65). (13 models ranked, as of 2026-10-05) - [Methodology](https://benchgecko.ai/gecko-tests/methodology): how the tests are built and graded. ### Latest findings - [DeepSeek V4 Flash 0731 identifies as another lab's model in 2 of 6 answers](https://benchgecko.ai/blog/finding-deepseek-v4-flash-0731-identifies-as-another-labs-model-v1): 2026-10-05 · Asked what it is and who made it, DeepSeek's DeepSeek V4 Flash 0731 named Google and Anthropic instead of DeepSeek in 2 of 6 replies (run 2026-10-04, no system prompt). - [DeepSeek V4.1 Flash identifies as another lab's model in 2 of 5 answers](https://benchgecko.ai/blog/finding-deepseek-v4-1-flash-identifies-as-another-labs-model-v1): 2026-10-05 · Asked what it is and who made it, DeepSeek's DeepSeek V4.1 Flash named OpenAI instead of DeepSeek in 2 of 5 replies (run 2026-10-04, no system prompt). - [Llama 4 Maverick identifies as another lab's model in 3 of 6 answers](https://benchgecko.ai/blog/finding-llama-4-maverick-identifies-as-another-labs-model-v1): 2026-10-04 · Asked what it is and who made it, Meta's Llama 4 Maverick named Google and OpenAI instead of Meta in 3 of 6 replies (run 2026-10-04, no system prompt). - [DeepSeek V4 Pro 0813 identifies as another lab's model in 3 of 6 answers](https://benchgecko.ai/blog/finding-deepseek-v4-pro-0813-identifies-as-another-labs-model-v1): 2026-10-04 · Asked what it is and who made it, DeepSeek's DeepSeek V4 Pro 0813 named OpenAI and Anthropic instead of DeepSeek in 3 of 6 replies (run 2026-10-04, no system prompt). ## Models and benchmarks - [All models](https://benchgecko.ai/models): leaderboard by BenchGecko score (normalized average of public benchmark scores), with price, context and license. - [Model pages](https://benchgecko.ai/model/claude-opus-5-5): /model/ · score, rank, list price, providers, benchmarks with sources, Gecko Tests grades, price history. - [Benchmarks](https://benchgecko.ai/benchmarks): every tracked benchmark and its ranked models. - [Compare](https://benchgecko.ai/compare): /compare/-vs- side by side. - [Methodology](https://benchgecko.ai/methodology): sources, normalization and refresh schedule. Top 10 by BenchGecko score (as of 2026-10-05): - 1. [GPT-5.5 Pro](https://benchgecko.ai/model/gpt-5-5-pro) (OpenAI): score 99.9, list price $30.00 input / $180.00 output per 1M tokens - 2. [Claude Mythos Preview](https://benchgecko.ai/model/claude-mythos-preview) (Anthropic): score 99.8, list price n/a input / n/a output per 1M tokens - 3. [Claude Opus 5.5](https://benchgecko.ai/model/claude-opus-5-5) (Anthropic): score 96.8, list price $4.00 input / $20.00 output per 1M tokens - 4. [GPT-6 Astra](https://benchgecko.ai/model/gpt-6-astra) (OpenAI): score 96.3, list price $10.00 input / $50.00 output per 1M tokens - 5. [DeepSeek V3.2 Speciale](https://benchgecko.ai/model/deepseek-v3-2-speciale) (DeepSeek): score 95.2, list price $0.40 input / $1.20 output per 1M tokens - 6. [Step 3.5 Flash](https://benchgecko.ai/model/step-3-5-flash) (stepfun): score 89.5, list price $0.10 input / $0.30 output per 1M tokens - 7. [Claude Sonnet 5.5](https://benchgecko.ai/model/claude-sonnet-5-5) (Anthropic): score 89.2, list price $2.00 input / $10.00 output per 1M tokens - 8. [GPT-5 Chat](https://benchgecko.ai/model/gpt-5-chat) (OpenAI): score 89, list price $1.25 input / $10.00 output per 1M tokens - 9. [Claude Fable 5.1](https://benchgecko.ai/model/claude-fable-5-1) (Anthropic): score 88.2, list price $10.00 input / $50.00 output per 1M tokens - 10. [Qwen2.5 72B Instruct Abliterated](https://benchgecko.ai/model/huihui-ai-qwen25-72b-instruct-abliterated) (HuiHui AI): score 87.2, list price n/a input / n/a output per 1M tokens ## Pricing - [AI pricing](https://benchgecko.ai/pricing): every model and every provider, filtered by use case and budget. - [Price per provider](https://benchgecko.ai/pricing/arbitrage): the same model at every provider, cheapest first; /pricing/arbitrage/ per model. - [Price drops](https://benchgecko.ai/pricing/drops): recorded price changes. - [Free tiers](https://benchgecko.ai/pricing/free): free and free-quota APIs. ## Companies, mindshare and infrastructure - [AI economy](https://benchgecko.ai/economy): valuations, funding, revenue and the AI Bubble Index; each figure has an as-of date and a source. - [Company pages](https://benchgecko.ai/economy/companies): /economy/company/. - [Mindshare](https://benchgecko.ai/mindshare): daily attention share from Hacker News, Wikipedia pageviews and GitHub stars. Top model mindshare (as of 2026-10-04): GPT-6 24.1% · Qwen 22.7% · GLM 14.2% · GPT-5 7.0% · Gemini 3 5.9% - [Hardware](https://benchgecko.ai/hardware): AI chips with specs and daily GPU rental prices (Vast.ai marketplace median per GPU hour). - [Providers](https://benchgecko.ai/providers): every API provider and its models. - [Status](https://benchgecko.ai/status): provider uptime pings. ## Optional - [Full context](https://benchgecko.ai/llms-full.txt): this file plus every Gecko Test result row, price spreads, company figures, GPU rental prices and the AI glossary. - [Glossary](https://benchgecko.ai/learn/glossary): AI terms with definitions linked to live data. - [Blog](https://benchgecko.ai/blog): weekly data roundups and findings.