PricingReading · ~3 min · 40 words deep

Tokenizer tax

The tokenizer tax is the extra cost of non-English text · tokenizers split many languages into more tokens than English, and models bill per token.

Text reviewed October 5, 2026

Tokenizer Tax test
TL;DR

The tokenizer tax is the extra cost of non-English text · tokenizers split many languages into more tokens than English, and models bill per token.

Level 1

Models do not read words, they read tokens. Tokenizers are usually trained mostly on English, so the same sentence in Japanese, Hindi or Arabic can take more tokens. Since prices are per token, the same request can cost more in another language, and also fills the context window faster.

Level 2

BenchGecko's Tokenizer Tax test sends the same text in several languages to each model and records how many input tokens the provider counts, relative to English. Differences between models come from their tokenizers: larger multilingual vocabularies usually lower the tax.

Level 3

The ratio depends on script and morphology, not only on vocabulary size. Because the count comes from the provider's own billing, it reflects what you actually pay, including any special tokens added by the serving stack.

The takeaway for you
If you are a
Curious · Normie
  • ·Some languages cost more to use with AI
If you are a
Builder
  • ·Budget non-English workloads with the measured ratio
  • ·Check /gecko-tests/tokenizer-tax before choosing a model
If you are a
Investor
  • ·A hidden price difference between markets
If you are a
Researcher
  • ·Token count ratios vs English per model
  • ·Driven by tokenizer vocabulary and training data mix
It depends on the model. The Tokenizer Tax test shows the measured ratio per language for every tested model.