Context · 32K+

Cheapest 32K context LLMs

Every LLM with at least 32K token context. Ranked by input price per 1M tokens.

Models50
Cheapest$-1000000.00
Min context32K tokens
What this page is
32K context is the floor for modern LLM use. This tier captures the broadest set of priced models at one of the deepest discounts. Ideal for chat, short-doc RAG, classification, and any workload where you do not need a huge window.

32K+ context models, cheapest first.

#ModelIn $/1MOut $/1MType
1openrouter logoAuto Router$-1000000.00$-1000000.00Closed
2openrouter logoBody Builder (beta)$-1000000.00$-1000000.00Closed
3openrouter logoFusion$-1000000.00$-1000000.00Closed
4openrouter logoPareto Code Router$-1000000.00$-1000000.00Closed
5openrouter logoElephant$0.00$0.00Closed
6openrouter logoFree Models Router$0.00$0.00Closed
7Google DeepMind logoGemma 3 12B (free)$0.00$0.00OSS
8Google DeepMind logoGemma 3 27B (free)$0.00$0.00OSS
9Google DeepMind logoGemma 3 4B (free)$0.00$0.00OSS
10Google DeepMind logoGemma 4 26B A4B (free)$0.00$0.00OSS
11Google DeepMind logoGemma 4 31B (free)$0.00$0.00OSS
12z-ai logoGLM 4.5 Air (free)$0.00$0.00OSS
13OpenAI logogpt-oss-120b (free)$0.00$0.00OSS
14OpenAI logogpt-oss-20b (free)$0.00$0.00OSS
15nousresearch logoHermes 3 405B Instruct (free)$0.00$0.00OSS
16tencent logoHy3 (free)$0.00$0.00Closed
17tencent logoHy3 preview (free)$0.00$0.00Closed
18
P
Laguna M.1 (free)
$0.00$0.00OSS
19
P
Laguna XS.2 (free)
$0.00$0.00OSS
20liquid logoLFM2.5-1.2B-Instruct (free)$0.00$0.00OSS
21liquid logoLFM2.5-1.2B-Thinking (free)$0.00$0.00OSS
22Meta logoLlama 3.2 3B Instruct (free)$0.00$0.00OSS
23Meta logoLlama 3.3 70B Instruct (free)$0.00$0.00OSS
24Meta logoLlama Guard 4 12B (free)$0.00$0.00Closed
25Google DeepMind logoLyria 3 Clip Preview$0.00$0.00Closed
26Google DeepMind logoLyria 3 Pro Preview$0.00$0.00Closed
27minimax logoMiniMax M2.5 (free)$0.00$0.00OSS
28Mistral AI logoMistral Small 3.1 24B (free)$0.00$0.00OSS
29NVIDIA logoNemotron 3 Nano 30B A3B (free)$0.00$0.00OSS
30NVIDIA logoNemotron 3 Nano Omni (free)$0.00$0.00OSS
31NVIDIA logoNemotron 3 Super (free)$0.00$0.00OSS
32NVIDIA logoNemotron 3 Ultra (free)$0.00$0.00OSS
33NVIDIA logoNemotron 3.5 Content Safety (free)$0.00$0.00OSS
34NVIDIA logoNemotron Nano 12B 2 VL (free)$0.00$0.00OSS
35NVIDIA logoNemotron Nano 9B V2 (free)$0.00$0.00OSS
36nex-agi logoNex-N2-Pro (free)$0.00$0.00OSS
37Cohere logoNorth Mini Code (free)$0.00$0.00OSS
38openrouter logoOwl Alpha$0.00$0.00Closed
39baidu logoQianfan-OCR-Fast (free)$0.00$0.00Closed
40Alibaba Qwen logoQwen3 4B (free)$0.00$0.00OSS
41Alibaba Qwen logoQwen3 Coder 480B A35B (free)$0.00$0.00OSS
42Alibaba Qwen logoQwen3 Next 80B A3B Instruct (free)$0.00$0.00OSS
43Alibaba Qwen logoQwen3.6 Plus (free)$0.00$0.00Closed
44Alibaba Qwen logoQwen3.6 Plus Preview (free)$0.00$0.00OSS
45stepfun logoStep 3.5 Flash (free)$0.00$0.00OSS
46arcee-ai logoTrinity Large Preview (free)$0.00$0.00OSS
47arcee-ai logoTrinity Mini (free)$0.00$0.00OSS
48cognitivecomputations logoUncensored (free)$0.00$0.00OSS
49liquid logoLFM2-2.6B$0.01$0.02OSS
50liquid logoLFM2-8B-A1B$0.01$0.02OSS
Cheapest
Auto Router
$-1000000.00/M
$ per 1M input tokens
Why the gap

At this tier, price reflects raw model quality more than context size. The cheapest 32K model is often a small open-source model; the most expensive is usually a frontier model deliberately billed the same across windows.

Most expensive
LFM2-2.6B
$0.01/M
$ per 1M input tokens
For most applications, yes. Chat histories, RAG retrievals, and single-document QA rarely exceed 16K tokens. 32K leaves plenty of headroom.