Devstral Review: Second-Fastest 9/9, Zero Reasoning Tokens (and Codestral Missed One)
Devstral 2 answered all nine of our coding tasks without emitting a single reasoning token, and it did it in a 2.1-second mean — the second fastest of the 75 models that scored 9 out of 9 as of 2026-10-02, a tenth of a second behind Solar Mini 4. That is the headline of this Devstral review. The footnote is more interesting: Mistral's other code model, Codestral 2508, was slower at 2.5 seconds and got one task wrong. Devstral costs $0.23 per 1,000 tasks at the list price captured 2026-10-02 — about 2.1 times Codestral's $0.11, and roughly 7.7 times Solar Mini 4's $0.03.
Mistral has more than one coding model on the same API. We ran two of them on the same day, on the same nine prompts. The newer, pricier one was faster, used fewer output tokens and got everything right.
The result
| Metric | Devstral 2 2512 | Codestral 2508 | Solar Mini 4 |
|---|---|---|---|
| Score | 9/9 | 8/9 | 9/9 |
| Missed | — | parse_csv_line | — |
| Derived cost / 1,000 tasks | $0.23 | $0.11 | $0.03 |
| Mean latency | 2.1s | 2.5s | 2s |
| Reasoning tokens per call | 0 | 0 | 0 |
| Tokens across the suite | 696 in / 911 out | 588 in / 950 out | 1,009 in / 938 out |
| Priced at (per 1M, in / out) | $0.4 / $2 | $0.3 / $0.9 | $0.05 / $0.2 |
| Context window | 262,144 | 256,000 | 524,288 |
| Rank among 75 at 9/9 (cost · speed) | 11th · 2nd | not ranked (8/9) | 1st · 1st |
All three were measured on 2026-10-02 and all three costs are derived at the list price captured that day. Codestral is not ranked: our cost and speed rankings only count models that solved every task, so its $0.11 (priced 2026-10-02) sits outside the comparison it would otherwise win against Devstral.
Devstral against Codestral, same vendor
The interesting comparison in this Devstral review is not with another lab. It is with Mistral itself.
On paper these two should be close. Both returned zero reasoning tokens. Their output volumes are almost identical — Devstral wrote 911 output tokens across the suite, Codestral 950, a difference of 39 tokens over nine answers. Neither model rambles.
The differences are three, and all three are measured:
- Correctness. Codestral got
parse_csv_linewrong — a wrong answer from executed code against hidden asserts, not an API failure. Devstral passed it. - Speed. Devstral's mean was 2.1 seconds against Codestral's 2.5, so the model sold for agentic work was 0.4 seconds faster per task than the one with code in its name.
- Price. Codestral was cheaper at $0.11 against $0.23, both priced 2026-10-02. Because the output token counts are so close, that gap is almost entirely the rate card: Devstral is priced at $2 per million output tokens, Codestral at $0.9, about 2.2 times more.
Devstral also consumed 696 input tokens on the same nine prompts where Codestral consumed 588 — 108 more, which points to a different tokenizer or chat template. It is small and it barely moves the bill, but it is a reminder that identical text is not identical input.
parse_csv_line is not a Codestral-only trap. In the same 2026-10-02 sweep, the only models to score 8/9 were Codestral 2508, Qwen3 Coder Plus and Qwen3 Coder Flash — and all three missed this one task, while plain Qwen3 Coder passed it. Earlier, DeepSeek V4-Pro lost its ninth point on the same function, as recorded in our AI coding ranking. Among fast, non-reasoning coding models, the CSV-line parser is where a 9/9 and an 8/9 part ways.
Speed without thinking
As of 2026-10-02 the three fastest models to clear the suite were Solar Mini 4 at 2 seconds, Devstral at 2.1 and GPT-5.4 mini at 2.3. All three report zero reasoning tokens. So does Qwen3 Coder, 6th fastest at 2.7 seconds. On tasks this size the fast tier is simply the tier that does not think before it answers, and Devstral sits squarely in it.
That is worth saying plainly because Devstral is not a small model — see below. A 0.1-second gap to Solar Mini 4 on a single scored run per task is not a meaningful difference. Treat the top three as one speed class.
Is it worth $0.23?
On this suite alone, no. Solar Mini 4 scored the same 9 out of 9 for $0.03, about 7.7 times less, a tenth of a second faster, with twice the context window. Qwen3 Coder scored 9/9 for $0.13. Both are priced 2026-10-02. Devstral is 11th cheapest of the 75 models that cleared the suite: inexpensive by any reasonable standard, not cheap by the standard of this sheet. For the low end of the field, our cheap coding roundup and the Ling 3.0 Flash VL review cover the alternatives.
Against the top of the market the picture flips. GPT-6 Astra scored the same 9/9 at $8.19 per 1,000 tasks, priced 2026-09-15. Nine self-contained Python functions cannot separate a frontier model from a competent small one, and they cannot separate Devstral from Solar Mini 4 either. What they can tell you is that Devstral does not fumble the easy work, and that it does it fast without burning reasoning tokens.
One more thing the per-task figure hides is shape. Our nine prompts are short and the answers are longer, so output tokens drive the bill — and Devstral's output rate of $2 per million is 5 times its input rate of $0.4. An agent loop is the opposite shape: it re-sends a repository and a growing history on every turn, so input dominates. There the number to compare is the input rate, priced 2026-10-02: $0.4 for Devstral, $0.3 for Codestral, $0.05 for Solar Mini 4 — 8 times less than Devstral. If you plan to run Devstral as an agent, estimate the bill from your own context size, not from our $0.23; our guide to cutting token costs in coding agents covers why input is where that money goes.
The case for Devstral rests on things we did not test: multi-file agentic editing, tool use and self-hosting. If those are why you are looking at it, our numbers are a floor check, not a verdict. If you want Mistral's general model instead, see the Mistral Medium 3.5 review.
What Devstral is, according to Mistral
The following is third-party and we have not verified it ourselves. Per Mistral's announcement post on mistral.ai (read 2026-10-02), Devstral 2 is an agentic coding model for autonomous software engineering, announced on 2025-12-09 alongside a smaller Devstral Small 2. The Hugging Face model card for Devstral-2-123B-Instruct-2512 (read 2026-10-02) describes it as a 123-billion-parameter dense transformer released under a modified MIT licence; the announcement gives Devstral Small 2 an Apache 2.0 licence. Read the modified licence yourself before assuming it is unrestricted. Mistral's own headline benchmark is SWE-bench Verified, which we did not run and do not reproduce here.
The endpoint we measured is mistralai/devstral-2512 on OpenRouter, the 123-billion-parameter model per the card above. The context window in our table is the figure OpenRouter lists.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; Codestral's miss was a wrong answer. Cost is derived from measured token counts multiplied by the list price captured 2026-10-02, not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
What we did not measure
- Agentic work, which is the entire pitch. We never gave Devstral a tool, a repository or a second turn. A model built to explore codebases was tested on nine standalone functions.
- SWE-bench or anything like it. We cite Mistral's positioning; we have no number of our own to set against it.
- The 262,144-token context. Our prompts are short.
- Self-hosted Devstral. Open weights run on your hardware will have different latency and no per-token price; our figures are for one hosted endpoint.
- Why Codestral missed
parse_csv_line. We recorded the wrong answer; we did not re-run it, so we cannot say whether it is a stable failure or a one-off at temperature 0. - The input-token gap. We did not inspect why Devstral counts 108 more input tokens than Codestral on identical prompts.
- Repeat runs. One scored attempt per task. A 0.1-second lead or deficit is inside that noise.
- Every coding specialist. KAT-Coder-Pro V2.5 returned HTTP 400 on every call on 2026-10-02, so it has no score and is not in any comparison here.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab