ConceptsReading · ~3 min · 41 words deep

Test-time compute

Test-time compute is the computation a model spends at answer time · spending more (longer reasoning, several attempts) can raise accuracy without retraining the model.

Text reviewed October 5, 2026

TL;DR

Test-time compute is the computation a model spends at answer time · spending more (longer reasoning, several attempts) can raise accuracy without retraining the model.

Level 1

Classic scaling made models better by training them bigger and longer. Reasoning models added a second lever: let the model think longer before it answers. More thinking tokens, or several attempts that are then compared, often improve results on math, code and planning.

Level 2

Common forms are long chain-of-thought reasoning, sampling several answers and voting, and search over candidate solutions with a verifier. The cost is paid on every request: more tokens, more latency and a higher bill, because reasoning tokens are billed as output.

Level 3

Gains from more test-time compute are largest on tasks with checkable answers and flatten out on others. Providers expose it as reasoning effort or thinking budget settings, so the same model can behave like several price and quality points.

The takeaway for you
If you are a
Curious · Normie
  • ·Letting the AI think longer before answering
If you are a
Builder
  • ·Use the lowest reasoning effort that meets your quality bar
  • ·Reasoning tokens are billed as output
If you are a
Investor
  • ·Inference demand grows with reasoning, not just with users
If you are a
Researcher
  • ·Second scaling axis next to training compute
  • ·Largest gains on verifiable tasks
Reasoning models are the main way to use it, but sampling several answers or running a search with a verifier also counts.