Test-time compute
Test-time compute is the computation a model spends at answer time · spending more (longer reasoning, several attempts) can raise accuracy without retraining the model.
Text reviewed October 5, 2026
Test-time compute is the computation a model spends at answer time · spending more (longer reasoning, several attempts) can raise accuracy without retraining the model.
Basic
Classic scaling made models better by training them bigger and longer. Reasoning models added a second lever: let the model think longer before it answers. More thinking tokens, or several attempts that are then compared, often improve results on math, code and planning.
Deep
Common forms are long chain-of-thought reasoning, sampling several answers and voting, and search over candidate solutions with a verifier. The cost is paid on every request: more tokens, more latency and a higher bill, because reasoning tokens are billed as output.
Expert
Gains from more test-time compute are largest on tasks with checkable answers and flatten out on others. Providers expose it as reasoning effort or thinking budget settings, so the same model can behave like several price and quality points.
Depending on why you're here
- ·Letting the AI think longer before answering
- ·Use the lowest reasoning effort that meets your quality bar
- ·Reasoning tokens are billed as output
- ·Inference demand grows with reasoning, not just with users
- ·Second scaling axis next to training compute
- ·Largest gains on verifiable tasks