Tokenizer tax
The tokenizer tax is the extra cost of non-English text · tokenizers split many languages into more tokens than English, and models bill per token.
Text reviewed October 5, 2026
The tokenizer tax is the extra cost of non-English text · tokenizers split many languages into more tokens than English, and models bill per token.
Basic
Models do not read words, they read tokens. Tokenizers are usually trained mostly on English, so the same sentence in Japanese, Hindi or Arabic can take more tokens. Since prices are per token, the same request can cost more in another language, and also fills the context window faster.
Deep
BenchGecko's Tokenizer Tax test sends the same text in several languages to each model and records how many input tokens the provider counts, relative to English. Differences between models come from their tokenizers: larger multilingual vocabularies usually lower the tax.
Expert
The ratio depends on script and morphology, not only on vocabulary size. Because the count comes from the provider's own billing, it reflects what you actually pay, including any special tokens added by the serving stack.
Depending on why you're here
- ·Some languages cost more to use with AI
- ·Budget non-English workloads with the measured ratio
- ·Check /gecko-tests/tokenizer-tax before choosing a model
- ·A hidden price difference between markets
- ·Token count ratios vs English per model
- ·Driven by tokenizer vocabulary and training data mix