Model Review

GPT-5.4 mini Review: 9/9 at $0.53 per 1,000 Tasks and 2.3 Seconds (Zero Reasoning Tokens)

We ran GPT-5.4 mini through our executed coding benchmark on 2026-07-30. It scored 9/9, at a measured $0.53 per 1,000 tasks, in 2.3 s average, with zero reasoning tokens. Two of those numbers matter together: 2.3 s is the fastest mean latency of any model that scored 9/9 among the 45 we have run as of 2026-08-06. Two models are faster or equal and neither is perfect — Llama 4 Scout at 1.5 s and Gemini 3 Flash Preview at 2.3 s both scored 8/9, each missing parse_csv_line. The next perfect scorer behind GPT-5.4 mini is Mistral Medium 3.5 at 2.9 s, then Claude Opus 5 Fast at 3.4 s, then GPT-5.4 at 3.6 s. On the same nine tasks, GPT-5.4 mini is 3.2x cheaper than GPT-5.4 ($0.53 against $1.69) and 16.7x cheaper than GPT-5.5 ($8.83) — all three scoring the same 9/9. The mechanism is the zero in the reasoning column, and it generalises further than this model does.

Horizontal bar chart of mean latency for six GPT-5-family models on our nine-task Python benchmark, GPT-5.4 mini fastest of the six at 2.3 seconds

Most model pages have one number worth reading. This one has two, and they only mean something together. GPT-5.4 mini averaged 2.3 s per task and scored 9/9 — the fastest mean latency of any model that scored 9/9 among the 45 we have run as of 2026-08-06. Every other model at the top of our latency table either scored 9/9 and was slower, or was fast and dropped a task. GPT-5.4 mini is the only model in the set holding both positions at once.

It is also cheap in a way the rate card does not fully explain. $0.53 per 1,000 completed tasks, against $1.69 for GPT-5.4 and $8.83 for GPT-5.5. Same nine tasks, same 9/9 from all three. That is 3.2x and 16.7x on identical results.

Read this before the tables. Our harness is nine short, self-contained Python functions — exactly the workload a small fast model handles well. It says nothing about long-context work, multi-file refactoring, agentic tool use, or ambiguous specs. Those are precisely where a mini tier separates from a flagship, and our harness cannot see any of them. Nothing on this page supports "always use the mini". It supports a narrower and more useful claim, which is that on bounded, clearly specified generation, the flagship premium bought nothing we could detect.

What we measured on 2026-07-30

One run, on 2026-07-30, priced at rates captured from the live catalog the same day:

The 9/9 on its own is not a distinguishing result and we will not sell it as one. 33 of the 45 models we have run scored 9/9, including everything from DeepSeek V3.2 at $0.08 per 1,000 tasks up to Gemini 3.1 Pro at $14.70 — a 184x cost spread across models that produced identical scores. Nine bounded Python functions sit inside the competent range of every serious 2026 coding model. Correctness is not the axis that separates them; cost and latency are.

What is distinguishing is where GPT-5.4 mini lands on those two axes at the same time.

The fastest 9/9 of the 45 models we have run

Here are the seven fastest of the 45 models we have run as of 2026-08-06, in order, with nothing omitted. Two of them are faster than or equal to GPT-5.4 mini and both dropped a task.

ModelScoreMean latencyMeasured cost / 1k tasksReasoning tokens / taskPriced at
Llama 4 Scout8/91.5 s$0.0302026-08-06
Gemini 3 Flash Preview8/92.3 s$0.3602026-08-06
GPT-5.4 mini9/92.3 s$0.5302026-07-30
Mistral Medium 3.59/92.9 s$0.8702026-07-17
Claude Opus 5 Fast9/93.4 s$10.2062026-07-30
GPT-5.49/93.6 s$1.6902026-07-29
Claude Haiku 4.59/93.7 s$0.9402026-07-29

Llama 4 Scout ran 1.5 s at $0.03 per 1,000 tasks — the lowest figure in both columns across the 45 models we have run as of 2026-08-06. It scored 8/9. Gemini 3 Flash Preview matched GPT-5.4 mini's 2.3 s at $0.36 and also scored 8/9. Both missed the same task, parse_csv_line, which is the most-missed task across our whole set. So the honest ranking is not "fastest"; it is fastest among the models that got everything right, and that claim survives the two models that beat GPT-5.4 mini on the clock.

Two more things fall out of the table. First, GPT-5.4 mini is 0.6 s ahead of the next perfect scorer and 1.1 s ahead of the one after that — a 26% gap to the second-fastest 9/9, which is wide for a leaderboard where the middle of the pack is separated by tenths. Second, six of the seven fastest emitted zero reasoning tokens. The one that did not, Claude Opus 5 Fast, emitted 6 and costs $10.20 per 1,000 tasks — 19.2x GPT-5.4 mini for a result that is 1.1 s slower and scored the same.

Five of those seven clear the bar on correctness. If your selection criterion is latency on bounded generation and you want a clean nine, the whole decision collapses into the cost column, and there GPT-5.4 mini is the cheapest of the five. If a dropped task in nine is acceptable, Llama 4 Scout is 0.8 s faster and 17.7x cheaper, and that trade is the one to weigh — the AI coding ranking sorts the full field on both axes.

Zero reasoning tokens is the whole mechanism

The useful part of this page is not GPT-5.4 mini. It is why it is fast, because that reason applies to models we have not run yet.

Compare it against GPT-5 mini — the other small OpenAI tier in our set. Both scored 9/9 on the same nine tasks. Then:

Two OpenAI mini tiers · both 9/9Reasoning tokens / taskMean latencyMeasured cost / 1k tasksList price in / out per 1MPriced at
GPT-5.4 mini02.3 s$0.53$0.75 / $4.502026-07-30
GPT-5 mini55515.2 s$1.53$0.25 / $22026-07-29

Same score, same tasks, same harness settings. One is 6.6x faster and 2.9x cheaper. The difference is the first column: GPT-5 mini spent 555 reasoning tokens per task thinking about function specs that GPT-5.4 mini answered without thinking at all.

Now look at the list-price column, because it is the opposite of the outcome. GPT-5 mini is the cheaper model on paper — $0.25 / $2 against $0.75 / $4.50, three times less on input and 2.25 times less on output. It still measured 2.9x more per completed task. Reasoning tokens bill at the output rate, so 555 of them per task at $2 per 1M works out to roughly $0.0011 of the $0.00153 that each GPT-5 mini task actually cost — about 73% of the bill, spent on tokens you never see. That figure is arithmetic from the stated rate and our measured token count, not a separate measurement, and it assumes reasoning bills at the output rate as OpenAI documents.

So the lesson is not the one the naming convention implies. Small does not mean fast. Not thinking means fast. A mini model with reasoning on is slower than a full-size model with reasoning off, and its per-token discount mostly evaporates. GPT-5 mini lists at $0.25 / $2 against GPT-5.4's $2.50 / $15 — 10x cheaper on input, 7.5x cheaper on output — and still finished at $1.53 against $1.69, under 10% cheaper per completed task, for 4.2x the latency (15.2 s against 3.6 s). A 10x rate-card advantage bought a rounding error. Against GPT-5.4 mini it loses outright on both counts, by a wide margin.

This pattern holds across the wider set. 14 of the 45 models we have run emitted zero reasoning tokens, and they cluster hard at the fast end: six of the seven fastest models are in that group, including both of the 8/9 models that beat or matched GPT-5.4 mini on the clock. At the other end, the slowest results in our data are reasoning-heavy — Gemini 2.5 Pro at 2,671 reasoning tokens per task took 25.3 s and cost $28.58 per 1,000 tasks, for 6/9. Reasoning tokens are the single best predictor of latency in our dataset, better than parameter count, tier name or price. If you are tuning for response time, that is the knob, and we go further into it in how to reduce LLM latency.

One honest qualifier: we sent no effort or reasoning parameter on any call. Zero reasoning tokens is what GPT-5.4 mini chose on nine easy specs at default settings, not a property we forced or a guarantee about harder prompts. A prompt that triggers thinking would move both the latency and the cost, and we have not measured that.

What nine Python functions cannot tell you

This section is here, in the middle, on purpose. A page that leads with "fastest and perfect" owes you the boundary of the claim before the recommendation, not after it.

Our nine tasks are short, self-contained Python functions with a clear prose spec and a fixed signature. two_sum, valid_parentheses, parse_csv_line and six like them. Each fits in a few hundred tokens. There is no ambiguity to resolve, no existing codebase to respect, no file to find, no tool to call, and no second turn. That is the exact shape of problem a small, fast model is good at.

The things a mini tier is expected to lose on are all outside the frame:

A 9/9 here is a floor test, not a ceiling test. It tells you GPT-5.4 mini will not fumble routine bounded generation. It tells you nothing about whether it holds up when the work gets hard, and the prior from every previous model generation is that mini tiers do not. Treat the measured numbers as a reason to route the easy majority of your calls to it, not as a reason to retire your flagship.

The whole measured GPT-5 ladder

Six GPT-5-family models make up the ladder below. Every one returned 9/9. Rows ordered by measured cost, cheapest first. Three further OpenAI results sit outside it and are listed after the chart.

ModelScoreMeasured cost / 1k tasksMean latencyReasoning tokens / taskList price in / out per 1MPriced at
GPT-5.4 mini9/9$0.532.3 s0$0.75 / $4.502026-07-30
GPT-5 mini9/9$1.5315.2 s555$0.25 / $22026-07-29
GPT-5.49/9$1.693.6 s0$2.50 / $152026-07-29
GPT-5.6 Sol9/9$4.986.6 s58$5 / $302026-07-29
GPT-5.59/9$8.8310.5 s176$5 / $302026-07-17
GPT-59/9$11.0022.9 s924$1.25 / $102026-07-30
The measured GPT-5 ladder: same 9/9, 2.3 s to 22.9 s per taskNine executed Python tasks, temperature 0, one scored attempt each. All six bars scored 9/9. Rows ordered by measured cost.GPT-5.4 mini2.3 sGPT-5 mini15.2 sGPT-5.43.6 sGPT-5.6 Sol6.6 sGPT-5.510.5 sGPT-522.9 sOne scale throughout: 22 px per second. Latency is the mean wall-clock time per task on our run, not a vendor figure.
Chart: DataLLM Lab. Mean latency per task, measured on our executed nine-task Python benchmark. All six models scored 9/9. Method: our methodology. Full run: the coding cost benchmark.

A 20.8x cost span and a 10.0x latency span, for identical results. $0.53 to $11.00, 2.3 s to 22.9 s, and the score column never moves. That is the single most useful fact in our dataset and it is not visible on any pricing page, because the pricing page prices tokens and this ladder prices completed work.

The pricing dates in that table are not uniform, and it matters. GPT-5.5 is priced at 2026-07-17. GPT-5 mini, GPT-5.4 and GPT-5.6 Sol are priced at 2026-07-29. GPT-5.4 mini and GPT-5 are priced at 2026-07-30. Each cost figure is the token count we measured multiplied by that model's list price on its stated date. Prices move — that is the reason we date them at all, and the reason we wrote LLM price volatility.

Two rows deserve a second look. GPT-5 is the most expensive result on the ladder despite listing at $1.25 / $10 — half GPT-5.4's input rate and a quarter of GPT-5.5's. It measured $11.00 because it spent 924 reasoning tokens per task. At $10 per 1M output, those reasoning tokens alone come to roughly $0.0092 of the $0.011 each task cost, about 84% — arithmetic from the stated rate, on the same assumption as before. And GPT-5.6 Sol and GPT-5.5 share an identical $5 / $30 list price and measured $4.98 against $8.83, a 1.8x gap at the same rate card, driven by token volume rather than price. We took that comparison apart in the GPT-5.6 Sol review.

Three more OpenAI results are not on that ladder. gpt-oss-120b, the open-weight release, scored 8/9 — it missed parse_csv_line — at a measured $0.05 per 1,000 tasks and 12.9 s, on 198 reasoning tokens per task, priced 2026-07-30. It is the second-cheapest result in our set of 45, behind Llama 4 Scout at $0.03, and the only OpenAI entry in the July runs that dropped a task, which is why it sits outside a table whose whole point is a flat score column. 10.6x cheaper than GPT-5.4 mini and 5.6x slower, with one function wrong: a different trade entirely, and not the one this page is about.

The other two ran on 2026-08-06, after the chart above was drawn, and are priced at that date. GPT-5.6 Terra Pro: 9/9, $6.99 per 1,000 tasks, 5 s, 246 reasoning tokens per task. GPT-5.6 Luna Pro: 8/9 — it also missed parse_csv_line — $3.79, 7.2 s, 548 reasoning tokens per task. Neither changes the argument on this page: Terra Pro matched GPT-5.4 mini's score at 13.2x the measured cost and 2.7 s more latency.

If you want the version of this argument that spans all four vendors rather than one, the AI coding ranking sorts the whole field, and GPT-5.4 vs GPT-5.5 takes the two flagship rungs head to head.

Four models scored 9/9 for less money

GPT-5.4 mini is the fastest 9/9 we have run. It is not the cheapest, and a page that only showed the GPT-5 ladder would leave you with that impression.

Four models beat it on measured cost with the same 9/9: DeepSeek V3.2 at $0.08, Qwen3 Coder Next at $0.10, DeepSeek Chat at $0.10, and DeepSeek V4-Flash at $0.13. That makes GPT-5.4 mini the fifth-cheapest perfect scorer in our set, not the first.

The interesting one is DeepSeek Chat. $0.10 per 1,000 tasks, 3.8 s, zero reasoning tokens, 9/9 — priced 2026-07-30. That is 5.3x cheaper than GPT-5.4 mini and only 1.5 s slower. If cost is what you are optimising and 3.8 s is inside your budget, GPT-5.4 mini is the wrong pick and we are not going to pretend otherwise. DeepSeek V3.2 at $0.08 is cheaper still, at 7.1 s.

The case for GPT-5.4 mini over those four is narrow and worth stating precisely: it is faster than all four — by 1.5 s against DeepSeek Chat and 12.2 s against DeepSeek V4-Flash — and it is on OpenAI infrastructure with OpenAI's ecosystem, availability and data-handling terms. Only the first of those is a benchmark result; the ecosystem, availability and data-handling terms are not something we measured. If the 1.5 s matters to a user staring at a spinner, $0.43 extra per 1,000 tasks is a rounding error and the choice is obvious. If the calls run in a batch queue overnight, it is 5.3x of pure waste. The cheap coding model roundup covers the sub-$1 tier in full.

Projected annually, at 1,000 tasks a day of roughly this size: about $193 a year on GPT-5.4 mini, against about $617 on GPT-5.4 and about $4,015 on GPT-5 — projections from our measured per-task figures at the stated rates, not bills anyone sent us. For your own token mix, the cost calculator does the arithmetic.

When to actually reach for GPT-5.4 mini

Use it as the default for bounded, clearly specified generation where a person is waiting. Autocomplete, single-function generation, formatting, extraction, classification, short transforms. On our nine tasks it scored the same 9/9 as GPT-5.5 at 16.7x less money and 8.2 s less latency per call. Nothing in our data justifies paying the flagship premium for work this shape.

Use it when latency is the product. 2.3 s is the fastest figure among the 33 models that scored 9/9, and the gap to the next perfect scorer is 0.6 s. Llama 4 Scout at 1.5 s and Gemini 3 Flash Preview at 2.3 s are faster or equal on the clock and each dropped a task, so if every one of nine has to be right, GPT-5.4 mini is the quickest way to get there. If you are inside a UI loop, that is the strongest argument on this page. If your calls run in a queue, speed is worth nothing and you should be reading the cost column instead, where four models beat it.

Do not use it as your only model. This is the recommendation the numbers do not support, and it is the one this page will be misread as making. Every capability where a mini tier is expected to fall behind a flagship — long context, multi-file work, tool use, ambiguity — is outside what we measured. Route the easy majority here and keep a flagship for the rest. Which ChatGPT model to use works through that routing decision task by task.

Do not read the mini label as a speed guarantee. GPT-5 mini is the counterexample sitting in our own data: same family, same "mini", 6.6x slower, 2.9x dearer, identical score. Check the reasoning-token behaviour of any model before you assume the small one is the fast one — the point we make at length in GPT-5 mini vs GPT-5 nano.

The general rule. Rank by cost and latency per completed task, not by per-token price or tier name. GPT-5 mini lists cheaper than GPT-5.4 mini and costs 2.9x more in practice. GPT-5 lists cheaper per token than GPT-5.4 and measured 6.5x more. Reasoning tokens are invisible on the rate card and they are frequently the majority of the bill.

Compare GPT-5.4 mini against GPT-5.4 on one key

One OpenAI-compatible endpoint, 300+ models, one key. Swap the model id between GPT-5.4 mini, GPT-5.4 and DeepSeek Chat, send your own prompts, and time them yourself.

How these numbers were produced

Nine executed Python tasks: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line.

The model receives a function signature and a prose spec. It never sees the assertions. The returned code runs against hidden assertions in an isolated python3 -I subprocess with a 12-second timeout. Every assertion passes or the task fails — no partial credit, no human grader, no LLM judge.

Temperature 0, max_tokens 4000, no effort or reasoning-budget parameter set. One scored attempt per task. The harness retries only on an API error, never on a wrong answer, which is why a miss would have stayed a miss.

Cost is derived, not invoiced. It is the token counts the API reported multiplied by that model's list price on a stated date. GPT-5.4 mini ran on 2026-07-30 and is priced at 2026-07-30 rates. Every figure here is true as of its pricing date and should be recomputed before you act on it.

Latency is wall-clock time per task on our run, from our machine, through one provider, at one moment. Network conditions and provider load move it. The ordering has been stable across our runs; the absolute seconds should not be treated as a vendor SLA.

Calls go through OpenRouter's OpenAI-compatible endpoint, deliberately not through the DataLLM Lab gateway. Nothing on this page depends on our infrastructure and you do not have to be our customer to reproduce it.

On the counting. Our original sweep was 13 models run in one sitting. GPT-5.4 mini is among the models run later on the same harness under the same settings, which brings the usable total to 45 as of 2026-08-06, of which 33 scored 9/9. Where this page says 45 models, that is the combined set. The core sweep was and remains 13.

Two entries in the run file are excluded from every count and every table. Claude Fable 5 returned empty responses with finish_reason=content_filter on four of the nine tasks after three retries each, so it has no comparable score and we will not quote one for it. Qwen3.5 397B-A17B returned an empty answer on token_bucket with finish_reason=length — it spent our 4,000-token ceiling on reasoning without emitting code. That ceiling is our harness constraint, not a defect in the model, and it is why the run is dropped rather than scored 8/9.

What we did not measure

Not measured at all: long-context work, multi-file refactoring, agentic or multi-turn tool use, any non-Python language, vision, and anything run locally. The harness is single-turn and API-hosted. It does not run agents and it does not pass tools.

The reasoning behaviour is default behaviour, not a property. Zero reasoning tokens is what GPT-5.4 mini did on nine easy specs with no effort parameter set. We did not test it at a raised effort level, and a prompt that triggers thinking would change both the latency and the cost. Read $0.53 and 2.3 s as a floor for this model, not a typical figure for all workloads.

A known scoring artefact: the 4,000-token ceiling can truncate a very verbose model mid-answer, and a truncated answer scores as a miss. That penalises verbosity, which is a real production cost but is not the same as being wrong. GPT-5.4 mini was not affected — it scored 9/9 — but the ceiling is part of the harness.

Nine tasks is nine data points. A 9/9 is a clean result on this set, not a claim about a distribution. We did not run it twice, we did not vary the prompts, and we did not buy any model a retry.

Never run by us, and never claimed as ours: GPT-5.4 nano, GPT-5 nano, Ollama or any locally-run model, and every retired Grok model. Run but excluded, with no comparable score: Claude Fable 5 and Qwen3.5 397B-A17B, for the reasons given above — both produced a partial denominator, so neither has a nine-task result. If a first-party nine-task score for any of these appears anywhere on this site, it is an error.

FAQ

Is GPT-5.4 mini actually as good as GPT-5.4?

On our nine executed Python tasks, yes — both scored 9/9, with GPT-5.4 mini at $0.53 per 1,000 tasks and 2.3 s against GPT-5.4's $1.69 and 3.6 s (priced 2026-07-30 and 2026-07-29). But those nine tasks are short, self-contained functions with unambiguous specs. They do not test long context, multi-file changes, tool use or ambiguity, which is where a mini tier is expected to fall behind. Our answer is "identical on this workload", not "identical".

Why is GPT-5.4 mini faster than GPT-5 mini when both are mini models?

Reasoning tokens. GPT-5.4 mini emitted zero per task; GPT-5 mini emitted 555. That is the whole difference — 2.3 s against 15.2 s, a 6.6x gap, at the same 9/9 score. Model size is not what set the latency here; whether the model chose to think before answering is. We sent no effort parameter to either.

GPT-5 mini has a lower list price. Why did it cost more?

Because reasoning tokens bill at the output rate. GPT-5 mini lists at $0.25 / $2 per 1M against GPT-5.4 mini's $0.75 / $4.50, so it is cheaper per token by roughly 3x on input. It still measured $1.53 per 1,000 tasks against $0.53 — 2.9x more — because it spent 555 reasoning tokens per task that never appear in the answer. At $2 per 1M that is around 73% of each task's cost, by arithmetic from the stated rate. Per-token price and cost-per-completed-task are different numbers.

Is GPT-5.4 mini the fastest or the cheapest model you have benchmarked?

Neither, outright. Llama 4 Scout is faster at 1.5 s and cheaper at $0.03 per 1,000 tasks, but it scored 8/9, and Gemini 3 Flash Preview matched 2.3 s at $0.36 with the same 8/9. Four models scored the full 9/9 for less money: DeepSeek V3.2 at $0.08, Qwen3 Coder Next at $0.10, DeepSeek Chat at $0.10 and DeepSeek V4-Flash at $0.13, which makes GPT-5.4 mini the fifth-cheapest perfect scorer in our set. The precise claim is narrower: at 2.3 s it is the fastest model that scored 9/9 among the 45 we have run as of 2026-08-06. DeepSeek Chat is the closest trade-off on price: 5.3x cheaper and 1.5 s slower.

Does a 9/9 mean GPT-5.4 mini can replace my flagship model?

No, and we would rather say so plainly. 33 of the 45 models we have run scored 9/9, so a perfect score on this harness is a floor test — it shows a model does not fumble routine bounded generation. It is silent on long context, multi-file refactoring, agentic tool use and ambiguous requirements, which is exactly where a mini tier separates from a flagship. Route the easy majority of calls to it and keep a flagship for the rest.

Is your measured cost the same as my bill?

No. We take the token counts the API reported and multiply by each model's list price on a stated date — 2026-07-30 for GPT-5.4 mini. It is a derived measured cost, not a vendor invoice. Your bill will differ with prompt length, caching, discounts, provider routing, effort level and any price change since that date. The ratios between models are the durable part; the absolute dollars are not. Full rate card in the GPT-5 API pricing guide.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.