PricingReading · ~3 min · 33 words deep

Batch API

A batch API takes a file of requests and returns the results later, usually within a day, at a discount to real-time pricing.

Text reviewed October 5, 2026

Model pricing
TL;DR

A batch API takes a file of requests and returns the results later, usually within a day, at a discount to real-time pricing.

Level 1

Not every AI job needs an instant answer. Evaluations, data labeling, summaries of archives and nightly reports can wait. Batch APIs let providers schedule that work when they have spare capacity, and they pass part of the saving on as a lower price.

Level 2

BenchGecko lists batch pricing as separate model entries marked "(batch)", so you can compare them directly with real-time prices. The trade-off is latency: results arrive within the provider's batch window, not in seconds.

Level 3

Batch discounts can combine with prompt caching on some providers. Rate limits for batch jobs are usually separate from real-time limits, which helps large offline workloads.

The takeaway for you
If you are a
Curious · Normie
  • ·A cheaper way to ask an AI many things when you can wait for the answers
If you are a
Builder
  • ·Move any job that can wait to batch
  • ·Compare (batch) entries on BenchGecko pricing pages
If you are a
Investor
  • ·Lets providers fill idle capacity
If you are a
Researcher
  • ·Offline evals are a natural fit
It depends on the provider and model. Compare the "(batch)" entries with the real-time entries on BenchGecko.