Model Reviews

Devstral Review: Second-Fastest 9/9, Zero Reasoning Tokens (and Codestral Missed One)

Devstral 2 answered all nine of our coding tasks without emitting a single reasoning token, and it did it in a 2.1-second mean — the second fastest of the 75 models that scored 9 out of 9 as of 2026-10-02, a tenth of a second behind Solar Mini 4. That is the headline of this Devstral review. The footnote is more interesting: Mistral's other code model, Codestral 2508, was slower at 2.5 seconds and got one task wrong. Devstral costs $0.23 per 1,000 tasks at the list price captured 2026-10-02 — about 2.1 times Codestral's $0.11, and roughly 7.7 times Solar Mini 4's $0.03.

DataLLM Lab article cover: Devstral Review: Second-Fastest 9/9, Zero Reasoning Tokens (and Codestral Missed One)

Mistral has more than one coding model on the same API. We ran two of them on the same day, on the same nine prompts. The newer, pricier one was faster, used fewer output tokens and got everything right.

The result

MetricDevstral 2 2512Codestral 2508Solar Mini 4
Score9/98/99/9
Missed—parse_csv_line—
Derived cost / 1,000 tasks$0.23$0.11$0.03
Mean latency2.1s2.5s2s
Reasoning tokens per call000
Tokens across the suite696 in / 911 out588 in / 950 out1,009 in / 938 out
Priced at (per 1M, in / out)$0.4 / $2$0.3 / $0.9$0.05 / $0.2
Context window262,144256,000524,288
Rank among 75 at 9/9 (cost · speed)11th · 2ndnot ranked (8/9)1st · 1st

All three were measured on 2026-10-02 and all three costs are derived at the list price captured that day. Codestral is not ranked: our cost and speed rankings only count models that solved every task, so its $0.11 (priced 2026-10-02) sits outside the comparison it would otherwise win against Devstral.

Devstral against Codestral, same vendor

The interesting comparison in this Devstral review is not with another lab. It is with Mistral itself.

On paper these two should be close. Both returned zero reasoning tokens. Their output volumes are almost identical — Devstral wrote 911 output tokens across the suite, Codestral 950, a difference of 39 tokens over nine answers. Neither model rambles.

The differences are three, and all three are measured:

Devstral also consumed 696 input tokens on the same nine prompts where Codestral consumed 588 — 108 more, which points to a different tokenizer or chat template. It is small and it barely moves the bill, but it is a reminder that identical text is not identical input.

parse_csv_line is not a Codestral-only trap. In the same 2026-10-02 sweep, the only models to score 8/9 were Codestral 2508, Qwen3 Coder Plus and Qwen3 Coder Flash — and all three missed this one task, while plain Qwen3 Coder passed it. Earlier, DeepSeek V4-Pro lost its ninth point on the same function, as recorded in our AI coding ranking. Among fast, non-reasoning coding models, the CSV-line parser is where a 9/9 and an 8/9 part ways.

Speed without thinking

Devstral: 2nd fastest of 75 at 9/9, 11th cheapestSame nine executed Python tasks. Outline bar = Codestral 2508, which scored 8/9 and is not ranked.MEAN LATENCY, SECONDSSolar Mini 42sDevstral 2 25122.1sGPT-5.4 mini2.3sCodestral 2508 · 8/92.5sQwen3 Coder2.7sDERIVED COST PER 1,000 TASKSSolar Mini 4$0.03Devstral 2 2512$0.23GPT-5.4 mini$0.53Codestral 2508 · 8/9$0.11Qwen3 Coder$0.13Latency scale: 150 px per second. Cost scale: 1,000 px per dollar. Every bar width = value x scale.Costs derived at list price on 2026-10-02, except GPT-5.4 mini, priced 2026-07-30.
The fast end of the 9/9 field is crowded with models that do not reason. Devstral is one of them.

As of 2026-10-02 the three fastest models to clear the suite were Solar Mini 4 at 2 seconds, Devstral at 2.1 and GPT-5.4 mini at 2.3. All three report zero reasoning tokens. So does Qwen3 Coder, 6th fastest at 2.7 seconds. On tasks this size the fast tier is simply the tier that does not think before it answers, and Devstral sits squarely in it.

That is worth saying plainly because Devstral is not a small model — see below. A 0.1-second gap to Solar Mini 4 on a single scored run per task is not a meaningful difference. Treat the top three as one speed class.

Is it worth $0.23?

On this suite alone, no. Solar Mini 4 scored the same 9 out of 9 for $0.03, about 7.7 times less, a tenth of a second faster, with twice the context window. Qwen3 Coder scored 9/9 for $0.13. Both are priced 2026-10-02. Devstral is 11th cheapest of the 75 models that cleared the suite: inexpensive by any reasonable standard, not cheap by the standard of this sheet. For the low end of the field, our cheap coding roundup and the Ling 3.0 Flash VL review cover the alternatives.

Against the top of the market the picture flips. GPT-6 Astra scored the same 9/9 at $8.19 per 1,000 tasks, priced 2026-09-15. Nine self-contained Python functions cannot separate a frontier model from a competent small one, and they cannot separate Devstral from Solar Mini 4 either. What they can tell you is that Devstral does not fumble the easy work, and that it does it fast without burning reasoning tokens.

One more thing the per-task figure hides is shape. Our nine prompts are short and the answers are longer, so output tokens drive the bill — and Devstral's output rate of $2 per million is 5 times its input rate of $0.4. An agent loop is the opposite shape: it re-sends a repository and a growing history on every turn, so input dominates. There the number to compare is the input rate, priced 2026-10-02: $0.4 for Devstral, $0.3 for Codestral, $0.05 for Solar Mini 4 — 8 times less than Devstral. If you plan to run Devstral as an agent, estimate the bill from your own context size, not from our $0.23; our guide to cutting token costs in coding agents covers why input is where that money goes.

The case for Devstral rests on things we did not test: multi-file agentic editing, tool use and self-hosting. If those are why you are looking at it, our numbers are a floor check, not a verdict. If you want Mistral's general model instead, see the Mistral Medium 3.5 review.

What Devstral is, according to Mistral

The following is third-party and we have not verified it ourselves. Per Mistral's announcement post on mistral.ai (read 2026-10-02), Devstral 2 is an agentic coding model for autonomous software engineering, announced on 2025-12-09 alongside a smaller Devstral Small 2. The Hugging Face model card for Devstral-2-123B-Instruct-2512 (read 2026-10-02) describes it as a 123-billion-parameter dense transformer released under a modified MIT licence; the announcement gives Devstral Small 2 an Apache 2.0 licence. Read the modified licence yourself before assuming it is unrestricted. Mistral's own headline benchmark is SWE-bench Verified, which we did not run and do not reproduce here.

The endpoint we measured is mistralai/devstral-2512 on OpenRouter, the 123-billion-parameter model per the card above. The context window in our table is the figure OpenRouter lists.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; Codestral's miss was a wrong answer. Cost is derived from measured token counts multiplied by the list price captured 2026-10-02, not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

What we did not measure

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.