Inception Mercury 2.5: 2.4 Seconds, 989 Reasoning Tokens, 17 Cents
Inception Mercury 2.5 scored 9 out of 9 on our executed Python benchmark with a 2.4-second mean latency, making it the 2nd fastest of the 56 models in our set that clear the suite — a hair behind GPT-5.4 Mini at 2.3s. The measured cost was $0.17 per 1,000 tasks, derived at list price on 2026-09-15. The number that does not fit is the one next to it: Mercury emitted 989 reasoning tokens per call. Every other model near the top of our speed ranking got there by not thinking. This one thinks about as hard as Gemini 3.8 Flash, which needed 10.2 seconds per call, and still answers in 2.4.
Ranking date: the 56-model ranks below describe the September 16, 2026 snapshot. In the October 2 set of 75 models at 9/9, Mercury 2.5 is fourth on mean latency and tenth on derived cost. The recorded 2.4 seconds and $0.17 have not changed.
In our data, reasoning tokens and wall-clock time move together so reliably that we mostly stopped checking. Mercury 2.5 is the first entry that breaks the relationship badly enough to be worth a page of its own.
The result
| Metric | Mercury 2.5 |
|---|---|
| Score | 9/9 |
| Mean latency | 2.4s — 2nd fastest of 56 at 9/9 |
| Measured cost / 1,000 tasks | $0.17 — 6th cheapest of 56 at 9/9 |
| Reasoning tokens per call | 989 |
| Tokens across the suite | 526 in / 9,882 out |
| List price in / out | $0.04 / $0.15 per 1M, captured 2026-09-15 |
| Context window | 260,000 |
| Measured | 2026-09-16 |
Sixth on cost and second on speed, out of 56 models that all answered every task correctly. Nothing else in our set holds both positions at once: the cheapest model, Ling 3.0 Flash VL at $0.07 (priced 2026-09-15), sits 10th on latency, and the fastest, GPT-5.4 Mini, sits 10th on cost.
The fast models do not think. This one does
Here is the comparison that matters, because the two models are separated by a tenth of a second and by everything else:
| GPT-5.4 Mini | Mercury 2.5 | |
|---|---|---|
| Speed rank of 56 at 9/9 | 1st | 2nd |
| Mean latency | 2.3s | 2.4s |
| Reasoning tokens per call | 0 | 989 |
| Output tokens across the suite | 969 | 9,882 |
| Measured cost / 1,000 tasks | $0.53 | $0.17 |
| Cost rank of 56 at 9/9 | 10th | 6th |
| Context window | 400,000 | 260,000 |
| Priced on | 2026-07-30 | 2026-09-15 |
GPT-5.4 Mini is fast the ordinary way: it emits zero reasoning tokens and 969 output tokens across the whole suite, and it is out of the door. Mercury emits 9,882 output tokens across the suite against GPT-5.4 Mini's 969 — an order of magnitude more text — and finishes 0.1 second later. That is the finding. Elsewhere in our set, more tokens means more seconds.
Two of those gaps divide exactly. Gemini 3.8 Flash emits fewer reasoning tokens than Mercury — 901 against 989 — and its mean is 10.2s, which is 4.25 times Mercury's 2.4s (2.4 × 4.25 = 10.2). GLM 5.3 Flash at 24.9s is 10.375 times Mercury's mean (2.4 × 10.375 = 24.9), and its measured cost of $0.34 per 1,000 tasks (priced 2026-08-31) is exactly double Mercury's $0.17 (priced 2026-09-15). Two different run dates and two different price snapshots, so read those as separate measurements rather than a head-to-head.
The vendor's explanation, which we did not test
Inception describes Mercury as a diffusion language model — one that refines a whole block of tokens in parallel over a number of denoising passes, instead of emitting them strictly one after another. If that is what is running behind the endpoint, a decoupling of token count from wall-clock time is exactly what you would predict.
We did not test that claim and cannot test it. It is third-party context, repeated here because it is the obvious candidate explanation, not because we verified it. Our harness sees an OpenAI-compatible endpoint: it sends a prompt, counts tokens, times the response, and runs the code. We have been burned before by inferring architecture from behaviour — what you measure through an API is the endpoint, including its serving stack, batching and any speculative decoding on top, not the weights. The honest statement is narrower and still useful: this endpoint returns an order of magnitude more tokens than GPT-5.4 Mini in roughly the same wall-clock time. Why is the vendor's story.
Where 17 cents sits
$0.17 per 1,000 tasks, derived at the list price captured 2026-09-15 ($0.04 in, $0.15 out per 1M), puts Mercury 6th cheapest of the 56 models at 9/9. That is remarkable given the token volume: it burns 9,882 output tokens across the suite and is still cheaper than GPT-5.4 Mini, which burns 969. A $0.15 output price absorbs a lot of verbosity.
It is not the floor. Ling 3.0 Flash VL holds that at $0.07 (priced 2026-09-15), with DeepSeek V3.2 at $0.08 (priced 2026-07-30) behind it — both with far fewer tokens and no need for a diffusion story. What Mercury offers over them is the latency column, and latency is the thing most cost tables ignore. If you are putting a model in front of a user, 2.4s versus 7.1s is a product decision; the gap from $0.08 to $0.17 per thousand calls is not.
What 2.4 seconds does not buy
- Difficulty. Nine self-contained Python functions cannot separate a frontier model from a competent small one. Fifty-six models scored 9/9. This suite measures whether a model can write a correct short function at all, and then how fast and how cheaply — it does not measure reasoning depth, and the 989 reasoning tokens Mercury spends here are spent on easy problems.
- Context. 260,000 tokens. Mid-field, and a quarter of what the million-token models carry.
- Consistency. One scored attempt per task. A 0.1-second gap between 1st and 2nd place on latency is not a ranking, it is a tie we happened to break.
- Long-horizon behaviour. A model that emits 989 reasoning tokens on a two-line function is a model whose token use on a long agent loop we cannot predict from this data.
For the wider field, the cheapest-API roundup and the full coding cost benchmark have the rest of the 56.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. The generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Runs go through OpenRouter. Cost is derived from measured token counts at the list price captured on the date shown next to each figure — it is not a billing statement, and list prices move.
An API-layer failure is recorded separately from a wrong answer, and a run with unscored tasks is marked excluded rather than scored. IBM Granite 4.2 8B is the current example: two of its nine tasks failed at the API layer after retries, so seven were scored and we publish no score for it. Seven correct out of seven attempted is not a 9/9 and we will not print it beside one. Its reasoning-token count — 1,483 per call, above the 1,327 of GLM 5.1, the heaviest reasoner among the 56 — is real, but it belongs to an incomplete run and is kept out of the chart above for that reason. We have also shipped a harness bug before that scored filtered empty responses as wrong answers, which is why anything short of full marks gets re-checked by hand. Full method on the methodology page.
What we did not measure
- Whether diffusion is what makes it fast. We measured an endpoint. We did not and could not inspect decoding.
- Time to first token. Our 2.4s is a mean of complete responses. For a diffusion decoder that distinction may matter more than usual — a model that refines a whole block in parallel has a very different streaming profile from one that emits left to right, and we have no streaming data at all.
- Latency variance. We report a mean. A 2.4s mean with an ugly tail would be a worse product than a 4s mean with none, and we cannot tell you which this is.
- The 260,000-token context. Our whole suite sends 526 input tokens.
- Anything harder than nine short functions. Including multi-file edits, tool use, and long agent loops.
- Repeat runs. One scored attempt per task, on 2026-09-16.
FAQ
How fast is Inception Mercury 2.5? 2.4 seconds mean on our suite, 2nd fastest of the 56 models that scored 9 out of 9. GPT-5.4 Mini is 1st at 2.3s.
What does Mercury 2.5 cost? $0.04 per million input tokens and $0.15 output, captured 2026-09-15, which worked out to a derived $0.17 per 1,000 tasks on our suite.
Is Mercury 2.5 a diffusion model? Inception says so. We did not verify it and our measurements cannot.
Did it get everything right? 9 out of 9, first attempt, on tasks that 56 models in our set also cleared. Treat the score as a floor check, not a ranking.
What is the catch? 989 reasoning tokens per call on trivial problems, and a 260,000-token context that is mid-field rather than large.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab