Pricing

Batch API Pricing: 70 Models Halve It, Nine Charge You More (Survey of 85)

Nine batch endpoints cost more than the standard endpoint they queue behind. Batch API pricing is supposed to be the easiest discount in the business — you surrender synchronous latency, you get half off. We captured the standard and :batch price pair for every base model publishing both on 2026-09-15. 85 base models have a :batch endpoint. 70 price it at exactly 50% of standard output. The other fifteen are the whole story, and the worst of them is openai/gpt-oss-120b: batch output at $0.60 per 1M against $0.17 standard, which the capture records as 353%. You would wait longer and pay more than triple.

DataLLM Lab article cover: Batch API Pricing: 70 Models Halve It, Nine Charge You More (Survey of 85)

Everyone treats the batch queue as free money: same model, no deadline, half the bill. That rule of thumb is right 70 times out of 85. The exceptions do not degrade gracefully — they invert.

What we captured

This is a catalogue survey, not a benchmark run. We took the published standard price and the published :batch price for every base model that offers both, as they stood on 2026-09-15, and compared the output rate — output is where batch discounts live and where nearly all of a coding or summarisation bill sits.

85 base models have a :batch endpoint. 70 price it at exactly 50% of standard output. Not approximately fifty. Exactly. That is a convention strong enough that most cost models hard-code it, which is precisely why the fifteen that ignore it are dangerous: nobody checks a constant.

Every endpoint that breaks the rule

All fifteen, sorted by batch output price as a share of standard output price. Prices are per 1M tokens, input first, captured 2026-09-15.

ModelStandard in / outBatch in / outBatch output vs standard
x-ai/grok-4.3$1.25 / $2.50$1.00 / $2.0080%
minimax/minimax-m3$0.30 / $1.20$0.30 / $1.20100%
thinkingmachines/inkling$1.00 / $4.05$1.00 / $4.05100%
z-ai/glm-5.3-flash$0.07 / $0.25$0.07 / $0.25100%
qwen/qwen3.8-2.4t-a95b$2.00 / $6.00$2.00 / $6.00100%
thinkingmachines/inkling-small$0.45 / $1.20$0.50 / $1.20100%
moonshotai/kimi-k3$2.65 / $13.28$3.00 / $15.00113%
deepseek/deepseek-v4-pro-0813$0.58 / $1.74$0.66 / $1.98114%
moonshotai/kimi-k2.7-code$0.71 / $3.50$0.95 / $4.00114%
nvidia/nemotron-3-ultra-550b-a55b$0.60 / $2.40$0.60 / $3.60150%
openai/gpt-oss-20b$0.03 / $0.13$0.05 / $0.20154%
qwen/qwen3.5-9b$0.10 / $0.15$0.17 / $0.25167%
google/gemma-4-31b-it$0.09 / $0.34$0.39 / $0.97285%
deepseek/deepseek-v4-flash-0731$0.06 / $0.11$0.11 / $0.33300%
openai/gpt-oss-120b$0.04 / $0.17$0.15 / $0.60353%

Six of these fifteen price batch output at or below standard output. The nine in bold price it higher — and one of the six, as the next section shows, still manages to cost more on the input side.

The nine that charge you to wait

Batch output price as a share of standard output priceNine endpoints price the batch queue above the synchronous one. Catalogue capture, 2026-09-15.50%100%openai/gpt-oss-120b353%deepseek/deepseek-v4-flash-0731300%google/gemma-4-31b-it285%qwen/qwen3.5-9b167%openai/gpt-oss-20b154%nvidia/nemotron-3-ultra-550b-a55b150%deepseek/deepseek-v4-pro-0813114%moonshotai/kimi-k2.7-code114%moonshotai/kimi-k3113%the other 70 models50%One scale throughout: 1.4 px per percentage point. Standard and batch prices captured 2026-09-15.
The dashed line at 100% is the break-even. Everything to its right is a batch endpoint that costs more than the synchronous one.

The extreme case is openai/gpt-oss-120b, an open-weight model whose whole appeal is being cheap. Standard output is $0.17 per 1M; batch output is $0.60 per 1M. Input moves the same direction, from $0.04 to $0.15. Both figures are from the 2026-09-15 capture. If a bulk-classification job routes to the batch queue because a config flag says batch is cheaper, this is the endpoint that quietly eats the saving and then some — which is one more argument for running gpt-oss-120b yourself. The smaller sibling behaves the same way: openai/gpt-oss-20b goes from $0.03 / $0.13 standard to $0.05 / $0.20 batch, recorded as 154%.

Three shapes of anomaly

The fifteen are not one phenomenon. They are three.

1. Partial discount (one endpoint). x-ai/grok-4.3 takes output from $2.50 to $2.00. That is a real discount, just a fifth rather than a half — $2.00 divided by $2.50 is 0.8 exactly, which is the 80% in the table. If you budgeted a 50% cut on Grok 4.3 batch work, you over-forecast the saving by a wide margin. Our Grok 4.3 review covers the standard endpoint.

2. No discount (five endpoints). minimax/minimax-m3, thinkingmachines/inkling, z-ai/glm-5.3-flash, qwen/qwen3.8-2.4t-a95b and thinkingmachines/inkling-small all publish a batch output price identical to standard. The batch endpoint exists; the discount does not. inkling-small is worse than a wash — output is unchanged at $1.20 while input rises from $0.45 to $0.50, so the batch endpoint is strictly more expensive than the synchronous one for every possible workload. Our reviews of Inkling and Inkling Small, and of MiniMax M3, all priced the standard endpoint.

3. Inverted (nine endpoints). The bars above. The mildest are the three around 113−114% — Kimi K3, deepseek-v4-pro-0813 and kimi-k2.7-code — where the premium is small enough to look like a rounding artefact. The steepest, gemma-4-31b-it at 285% and deepseek-v4-flash-0731 at 300%, are not artefacts of anything.

We do not know why any of these are priced this way, and we are not going to guess. A catalogue row can be a deliberate tier, a quirk of whoever serves that endpoint, or a stale field nobody has touched. What we can say is what the catalogue said on 2026-09-15: a cost model that assumes 50% is wrong on fifteen of eighty-five names, and prices move.

The one anomaly inside our benchmark set

Fourteen of the fifteen are models we have never put through our executed-Python suite. One is: z-ai/glm-5.3-flash, which scored 9 of 9.

Metricz-ai/glm-5.3-flash
Score9/9
Derived cost / 1,000 tasks$0.34 (list price captured 2026-08-31)
Mean latency24.9s
Reasoning tokens per call1,212
Tokens across the suite655 in / 11,884 out
Standard list at benchmark time$0.075 in / $0.25 out per 1M (2026-08-31)
Standard list in the batch capture$0.07 in / $0.25 out per 1M (2026-09-15)
Batch list$0.07 in / $0.25 out per 1M (2026-09-15)
Batch output vs standard100%

Two things fall out of that table. First, the half-off that would normally apply to the output side of that $0.34 does not exist on this model — and with 11,884 output tokens against 655 input tokens across the suite, output is where its bill lives. A GLM 5.3 Flash workload moved to the batch queue trades latency for nothing.

Second, and we are publishing this rather than smoothing it over: the two captures disagree about the standard input price. The benchmark's own capture on 2026-08-31 recorded $0.075 per 1M in; the batch-pair capture on 2026-09-15 recorded $0.07. Output identical at $0.25 in both, input field moved. We print both with their dates instead of picking a winner. Our GLM 5.3 Flash review uses the 2026-08-31 figure, which is the one $0.34 was derived from.

Worth keeping in proportion: 9 of 9 is also what GPT-6 Astra scored at $8.19 per 1,000 tasks and Claude Fable 5.1 at $8.09. Nine self-contained Python functions cannot separate a frontier model from a competent small one. They separate nothing at the top of the field. What they do well is price identical work, which is the only claim we are making here.

The latest-alias trap

Several of the anomalous ids are what a floating -latest alias resolves to, so you can land on one without ever typing its name. From our alias capture on 2026-09-15: ~deepseek/deepseek-v4-flash-latest resolves to deepseek/deepseek-v4-flash-0731 (300%), ~deepseek/deepseek-pro-latest to deepseek/deepseek-v4-pro-0813 (114%), ~moonshotai/kimi-latest to moonshotai/kimi-k3 (113%), and ~z-ai/glm-flash-latest to z-ai/glm-5.3-flash (100%). The full map is in our piece on what the latest aliases actually point at.

That is the compounding failure: a floating alias picks the model, a config constant picks the queue, and the assumed 50% is never checked against either. Estimate from measured token counts, and pin the exact id.

How these numbers were produced

The batch pairs are a catalogue capture, not a measurement. We recorded the published standard and :batch prices for every base model offering both, on 2026-09-15, and compared the output rates. No inference was run against any :batch endpoint. Every score and latency on this site comes from the standard synchronous endpoint.

The benchmark numbers come from nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task — the harness reconnects only when the API itself errors, never after a wrong answer. Cost is derived from measured token counts at the list price captured on the stated date, not a billing statement. Runs go through OpenRouter, deliberately: the numbers do not depend on our own infrastructure and you do not have to be our customer to reproduce them. Full method on the methodology page.

What we did not measure

For the standard-endpoint view of the cheap end, see the cheapest LLM APIs and how to cut API costs.

FAQ

Is batch API pricing always 50% off? No. Of 85 base models with a :batch endpoint, 70 price output at exactly 50% of standard and 15 do not, as of 2026-09-15.

Which batch endpoint is the worst value? openai/gpt-oss-120b, at 353% of its standard output price — $0.60 per 1M batch against $0.17 standard on 2026-09-15.

Are there batch endpoints with no discount? Five sit at exactly 100%: minimax-m3, inkling, glm-5.3-flash, qwen3.8-2.4t-a95b and inkling-small. The last is worse than neutral, since its input rises from $0.45 to $0.50.

Does the discount apply to input tokens too? Not reliably. We compared output rates; the anomalies show input moving independently in both directions.

Did you benchmark the batch endpoints? No. Every score on this site is from the standard synchronous endpoint. The batch figures are a dated price capture.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.