Batch API
A batch API takes a file of requests and returns the results later, usually within a day, at a discount to real-time pricing.
Text reviewed October 5, 2026
A batch API takes a file of requests and returns the results later, usually within a day, at a discount to real-time pricing.
Basic
Not every AI job needs an instant answer. Evaluations, data labeling, summaries of archives and nightly reports can wait. Batch APIs let providers schedule that work when they have spare capacity, and they pass part of the saving on as a lower price.
Deep
BenchGecko lists batch pricing as separate model entries marked "(batch)", so you can compare them directly with real-time prices. The trade-off is latency: results arrive within the provider's batch window, not in seconds.
Expert
Batch discounts can combine with prompt caching on some providers. Rate limits for batch jobs are usually separate from real-time limits, which helps large offline workloads.
Depending on why you're here
- ·A cheaper way to ask an AI many things when you can wait for the answers
- ·Move any job that can wait to batch
- ·Compare (batch) entries on BenchGecko pricing pages
- ·Lets providers fill idle capacity
- ·Offline evals are a natural fit