Model Reviews

Inception Mercury 2.5: 2.4 Seconds, 989 Reasoning Tokens, 17 Cents

Inception Mercury 2.5 scored 9 out of 9 on our executed Python benchmark with a 2.4-second mean latency, making it the 2nd fastest of the 56 models in our set that clear the suite — a hair behind GPT-5.4 Mini at 2.3s. The measured cost was $0.17 per 1,000 tasks, derived at list price on 2026-09-15. The number that does not fit is the one next to it: Mercury emitted 989 reasoning tokens per call. Every other model near the top of our speed ranking got there by not thinking. This one thinks about as hard as Gemini 3.8 Flash, which needed 10.2 seconds per call, and still answers in 2.4.

Ranking date: the 56-model ranks below describe the September 16, 2026 snapshot. In the October 2 set of 75 models at 9/9, Mercury 2.5 is fourth on mean latency and tenth on derived cost. The recorded 2.4 seconds and $0.17 have not changed.

DataLLM Lab article cover: Inception Mercury 2.5: 2.4 Seconds, 989 Reasoning Tokens, 17 Cents

In our data, reasoning tokens and wall-clock time move together so reliably that we mostly stopped checking. Mercury 2.5 is the first entry that breaks the relationship badly enough to be worth a page of its own.

The result

MetricMercury 2.5
Score9/9
Mean latency2.4s — 2nd fastest of 56 at 9/9
Measured cost / 1,000 tasks$0.17 — 6th cheapest of 56 at 9/9
Reasoning tokens per call989
Tokens across the suite526 in / 9,882 out
List price in / out$0.04 / $0.15 per 1M, captured 2026-09-15
Context window260,000
Measured2026-09-16

Sixth on cost and second on speed, out of 56 models that all answered every task correctly. Nothing else in our set holds both positions at once: the cheapest model, Ling 3.0 Flash VL at $0.07 (priced 2026-09-15), sits 10th on latency, and the fastest, GPT-5.4 Mini, sits 10th on cost.

The fast models do not think. This one does

Here is the comparison that matters, because the two models are separated by a tenth of a second and by everything else:

GPT-5.4 MiniMercury 2.5
Speed rank of 56 at 9/91st2nd
Mean latency2.3s2.4s
Reasoning tokens per call0989
Output tokens across the suite9699,882
Measured cost / 1,000 tasks$0.53$0.17
Cost rank of 56 at 9/910th6th
Context window400,000260,000
Priced on2026-07-302026-09-15

GPT-5.4 Mini is fast the ordinary way: it emits zero reasoning tokens and 969 output tokens across the whole suite, and it is out of the door. Mercury emits 9,882 output tokens across the suite against GPT-5.4 Mini's 969 — an order of magnitude more text — and finishes 0.1 second later. That is the finding. Elsewhere in our set, more tokens means more seconds.

Mean latency of five heavy-reasoning 9/9 modelsEvery model here scored 9/9 and emits at least 900 reasoning tokens per call. One of them is not slow.rtok / callMercury 2.59892.4sGemini 3.8 Flash90110.2sGLM 5.11,32723.6sGLM 5.3 Flash1,21224.9sQwen3.7 Max1,23625.8sOne scale throughout: 20 px per second of mean latency. Each model measured on its own run date.
Five models that all reason heavily and all score 9/9. Four of them pay for it in seconds.

Two of those gaps divide exactly. Gemini 3.8 Flash emits fewer reasoning tokens than Mercury — 901 against 989 — and its mean is 10.2s, which is 4.25 times Mercury's 2.4s (2.4 × 4.25 = 10.2). GLM 5.3 Flash at 24.9s is 10.375 times Mercury's mean (2.4 × 10.375 = 24.9), and its measured cost of $0.34 per 1,000 tasks (priced 2026-08-31) is exactly double Mercury's $0.17 (priced 2026-09-15). Two different run dates and two different price snapshots, so read those as separate measurements rather than a head-to-head.

The vendor's explanation, which we did not test

Inception describes Mercury as a diffusion language model — one that refines a whole block of tokens in parallel over a number of denoising passes, instead of emitting them strictly one after another. If that is what is running behind the endpoint, a decoupling of token count from wall-clock time is exactly what you would predict.

We did not test that claim and cannot test it. It is third-party context, repeated here because it is the obvious candidate explanation, not because we verified it. Our harness sees an OpenAI-compatible endpoint: it sends a prompt, counts tokens, times the response, and runs the code. We have been burned before by inferring architecture from behaviour — what you measure through an API is the endpoint, including its serving stack, batching and any speculative decoding on top, not the weights. The honest statement is narrower and still useful: this endpoint returns an order of magnitude more tokens than GPT-5.4 Mini in roughly the same wall-clock time. Why is the vendor's story.

Where 17 cents sits

$0.17 per 1,000 tasks, derived at the list price captured 2026-09-15 ($0.04 in, $0.15 out per 1M), puts Mercury 6th cheapest of the 56 models at 9/9. That is remarkable given the token volume: it burns 9,882 output tokens across the suite and is still cheaper than GPT-5.4 Mini, which burns 969. A $0.15 output price absorbs a lot of verbosity.

It is not the floor. Ling 3.0 Flash VL holds that at $0.07 (priced 2026-09-15), with DeepSeek V3.2 at $0.08 (priced 2026-07-30) behind it — both with far fewer tokens and no need for a diffusion story. What Mercury offers over them is the latency column, and latency is the thing most cost tables ignore. If you are putting a model in front of a user, 2.4s versus 7.1s is a product decision; the gap from $0.08 to $0.17 per thousand calls is not.

What 2.4 seconds does not buy

For the wider field, the cheapest-API roundup and the full coding cost benchmark have the rest of the 56.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. The generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Runs go through OpenRouter. Cost is derived from measured token counts at the list price captured on the date shown next to each figure — it is not a billing statement, and list prices move.

An API-layer failure is recorded separately from a wrong answer, and a run with unscored tasks is marked excluded rather than scored. IBM Granite 4.2 8B is the current example: two of its nine tasks failed at the API layer after retries, so seven were scored and we publish no score for it. Seven correct out of seven attempted is not a 9/9 and we will not print it beside one. Its reasoning-token count — 1,483 per call, above the 1,327 of GLM 5.1, the heaviest reasoner among the 56 — is real, but it belongs to an incomplete run and is kept out of the chart above for that reason. We have also shipped a harness bug before that scored filtered empty responses as wrong answers, which is why anything short of full marks gets re-checked by hand. Full method on the methodology page.

What we did not measure

FAQ

How fast is Inception Mercury 2.5? 2.4 seconds mean on our suite, 2nd fastest of the 56 models that scored 9 out of 9. GPT-5.4 Mini is 1st at 2.3s.

What does Mercury 2.5 cost? $0.04 per million input tokens and $0.15 output, captured 2026-09-15, which worked out to a derived $0.17 per 1,000 tasks on our suite.

Is Mercury 2.5 a diffusion model? Inception says so. We did not verify it and our measurements cannot.

Did it get everything right? 9 out of 9, first attempt, on tasks that 56 models in our set also cleared. Treat the score as a floor check, not a ranking.

What is the catch? 989 reasoning tokens per call on trivial problems, and a 260,000-token context that is mid-field rather than large.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.