MiMo v2.6 Review: Three Endpoints, Three 9/9s, and a 16.6x Speed Tax
Xiaomi's MiMo v2.6 ships as three endpoints, and all three scored 9 out of 9 on our executed Python benchmark on 2026-10-02. The cheapest, MiMo v2.6 Flash, cost $0.04 per 1,000 tasks and emitted 9 reasoning tokens per call. The fastest, MiMo v2.6 Pro UltraSpeed, answered in a 2.5-second mean and cost $4.80 — against $0.29 and 9.5 seconds for plain MiMo v2.6 Pro. That is 16.6 times the Pro price per task to save 7 seconds, on a variant OpenRouter describes as built from the same checkpoint. Same score, three bills, and the middle one is the hardest to justify.
Most model families give you a quality ladder: a small one, a big one, and a reason to pay for the big one. On nine self-contained Python functions MiMo v2.6 gives you no quality ladder at all — every rung scored 9 out of 9 — so what is left to compare is price and wall-clock time. Price alone spreads across a factor of 120, from $0.04 to $4.80 per 1,000 tasks.
This is a review of the models, not of Xiaomi's coding agent. If you came here for the CLI, that is a different product, covered in our MiMo-Code review.
The result
| Metric | v2.6 Flash | v2.6 Pro | v2.6 Pro UltraSpeed | v2.5 Pro (reference) |
|---|---|---|---|---|
| Score | 9/9 | 9/9 | 9/9 | 9/9 |
| Measured cost / 1,000 tasks | $0.04 | $0.29 | $4.80 | $0.67 |
| Mean latency | 5s | 9.5s | 2.5s | 26.1s |
| Reasoning tokens per call | 9 | 169 | 377 | 554 |
| Tokens across the suite (in / out) | 610 / 1,022 | 610 / 2,745 | 610 / 4,661 | 2,116 / 5,895 |
| Priced at, per 1M (in / out) | $0.14 / $0.28 | $0.435 / $0.87 | $4.35 / $8.7 | $0.435 / $0.87 |
| Price captured | 2026-10-02 | 2026-10-02 | 2026-10-02 | 2026-07-30 |
| Context window | 1,050,000 | 1,050,000 | 1,048,576 | 1,050,000 |
| Rank of 75 at 9/9: cost / speed | 2 / 25 | 12 / 46 | 55 / 5 | 19 / 74 |
The ranks are positions among the 75 models in our set that had scored 9 out of 9 as of 2026-10-02. All three v2.6 endpoints were run in the same sweep on the same day, on identical prompts — note the identical 610 input tokens.
UltraSpeed: what 7 seconds costs
Per OpenRouter's model listing (third-party, read 2026-10-02), Pro UltraSpeed is built from the same MiMo v2.6 Pro checkpoint, matches it in quality, and delivers much higher output speed. Its rate card, priced 2026-10-02, is $4.35 in and $8.7 out per million tokens, against $0.435 and $0.87 for Pro — exactly 10x on both sides.
The measured gap is wider than the rate card, because the two endpoints did not behave the same on identical prompts:
- Output: UltraSpeed wrote 4,661 tokens across the suite against Pro's 2,745 — about 1.7x as much.
- Reasoning: 377 reasoning tokens per call against 169 — about 2.2x.
- Cost: a 10x rate card applied to more tokens lands at $4.80 against $0.29, or 16.6x per task.
- Latency: a 2.5-second mean against 9.5 seconds. The saving is 7 seconds, which is a 3.8x wall-clock speed-up — a smaller gain than the 10x price step, partly because UltraSpeed spends some of its extra throughput writing more.
“Same checkpoint” and “same behaviour” are different claims. We have been caught by this before — the Ox Alpha episode taught us that what a model does is set by the endpoint serving it, not only by the weights. Here the score held, but the token footprint did not.
Is the speed worth it? UltraSpeed is the 5th fastest of the 75 models that scored 9/9 as of 2026-10-02, and it is 55th on cost. If a human is waiting on every call, 7 seconds per call is real. But against MiMo v2.6 Flash the trade is starker: Flash answered in 5 seconds for $0.04, so UltraSpeed costs 120 times Flash per task to save 2.5 seconds. And Solar Mini 4 scored the same 9 out of 9 in 2 seconds for $0.03 — faster and cheaper than every MiMo endpoint on this suite, as of 2026-10-02. For more ways to buy latency without buying a premium tier, see our latency guide.
Flash and its 9 reasoning tokens
On our suite Flash beat Pro on every axis we measure except score, where all three endpoints tie. Against Pro it was 7.25x cheaper per task — more than the 3.1x gap in the rate card, because it wrote about 2.7x fewer output tokens (1,022 against 2,745) — and it was faster, 5 seconds against 9.5. It emitted 9 reasoning tokens per call where Pro emitted 169. On problems of this size Flash simply decides not to think, and the asserts say it did not need to.
Reasoning tokens are billed as output, and they are usually what makes a cheap-looking model expensive; we tracked that pattern across dozens of models in the agent cost study. Flash is the opposite case: a low rate card and a short answer.
Against MiMo v2.5 Pro, on the same rate card
MiMo v2.5 Pro was measured on 2026-07-30 at $0.435 in and $0.87 out per million — the same rate card MiMo v2.6 Pro carries on 2026-10-02. So the difference between the two is entirely behaviour:
| v2.5 Pro | v2.6 Pro | |
|---|---|---|
| Measured cost / 1,000 tasks | $0.67 (priced 2026-07-30) | $0.29 (priced 2026-10-02) |
| Mean latency | 26.1s | 9.5s |
| Reasoning tokens per call | 554 | 169 |
| Output tokens across the suite | 5,895 | 2,745 |
| Input tokens across the suite | 2,116 | 610 |
Same price, same score, 2.3x cheaper per task and 16.6 seconds faster. Xiaomi did not cut the price; the model stopped over-thinking, with reasoning tokens per call falling from 554 to 169. As of 2026-10-02 v2.5 Pro sits 74th of 75 on speed among 9/9 models; v2.6 Pro sits 46th.
The input row is the one we cannot explain. The nine prompts are identical, yet v2.5 Pro recorded 2,116 input tokens against 610 for every v2.6 endpoint. A gap like that is usually provider-side text added before the user's prompt — the pattern we documented in hidden system prompt tokens — but we did not probe v2.5 Pro with a single request to confirm it, so treat it as an observation, not a diagnosis.
Second cheapest, not the floor
Before this sweep the cheapest model to clear our suite was Ling 3.0 Flash VL at $0.07, priced 2026-09-15. MiMo v2.6 Flash at $0.04 beats that. It is not the new floor, because Solar Mini 4 was measured in the same 2026-10-02 sweep at $0.03.
| Rank of 75 at 9/9 (cost) | Model | Cost / 1,000 tasks | Priced on |
|---|---|---|---|
| 1 | Solar Mini 4 | $0.03 | 2026-10-02 |
| 2 | MiMo v2.6 Flash | $0.04 | 2026-10-02 |
| 3 | Ling 3.0 Flash VL | $0.07 | 2026-09-15 |
| 4 | DeepSeek V3.2 | $0.08 | 2026-07-30 |
So: MiMo v2.6 Flash is the second-cheapest model to score 9/9 in our set as of 2026-10-02. One cent separates it from first place on a single scored run per task; a re-run could reorder those two. Of the top two, Solar Mini 4 is also faster (2 seconds against 5), while MiMo v2.6 Flash carries a 1,050,000-token context against Solar Mini 4's 524,288. The wider cheap end is in our cheap coding roundup.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; none of the three MiMo v2.6 endpoints had one. Cost is derived — measured input and output token counts multiplied by the list price captured on the run date, 2026-10-02 for the v2.6 endpoints — not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
What we did not measure
- Whether Pro is better than Flash at anything. Nine self-contained Python functions cannot separate a flagship from a competent small model. Every MiMo endpoint maxed the suite, so the suite has nothing left to say about quality between them. Pro's pitch is agentic and multimodal work; we tested neither.
- Images, audio and video. Third-party coverage describes both models as omnimodal. Our suite is text-only.
- The million-token context. Our prompts are a few dozen tokens each.
- Throughput. We record mean latency per call, not tokens per second, so we cannot check OpenRouter's output-speed claim directly — only that the wall-clock gain on our calls was 3.8x.
- Why UltraSpeed writes more. Same checkpoint per OpenRouter, 1.7x the output on our prompts. We observed it; we cannot explain it from outside.
- The v2.5 Pro input gap. We did not send it a single raw request, so the source of its extra 1,506 input tokens is unconfirmed.
- Repeat runs. One scored attempt per task. Single-run figures, not averages, which is why we call a one-cent gap a near tie.
Third-party sources: OpenRouter model listings for xiaomi/mimo-v2.6-flash, xiaomi/mimo-v2.6-pro and xiaomi/mimo-v2.6-pro-ultraspeed, and DataNorth's MiMo v2.6 release report, all read 2026-10-02.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab