Model Reviews

Codestral Review (25.08): Eight of Nine, Beaten by Its Own Sibling

This Codestral review comes down to one task. On our executed Python benchmark Codestral 25.08 scored 8 out of 9 at $0.11 per 1,000 tasks (priced 2026-10-02), with a 2.5-second mean and zero reasoning tokens. The run was complete — all nine tasks returned code — so the 8/9 is a real score, not an API accident. The one it got wrong was parse_csv_line, the same task that tripped Qwen3 Coder Plus and Qwen3 Coder Flash in the same sweep. Meanwhile Devstral 2, from the same vendor, solved all nine at $0.23 and was faster, at 2.1 seconds. Codestral is not the Mistral model to pick for writing whole functions from a spec, and Mistral, to be fair, never said it was.

DataLLM Lab article cover: Codestral Review (25.08): Eight of Nine, Beaten by Its Own Sibling

A model with “code” in its name is supposed to win a coding benchmark. On ours, the dedicated code-completion model from Mistral finished one task behind Mistral's agentic coding model, and that gap tells you something about what each was trained to do.

The result

MetricCodestral 25.08Devstral 2
Score8/99/9
Task missedparse_csv_line (wrong answer)none
Measured cost / 1,000 tasks (priced 2026-10-02)$0.11$0.23
Mean latency2.5s2.1s
Reasoning tokens per call00
Tokens across the suite588 in / 950 out696 in / 911 out
List price in / out, per 1M (2026-10-02)$0.3 / $0.9$0.4 / $2
Context window256,000262,144
Measured2026-10-022026-10-02

Both models were measured on the same day against the same nine prompts. Neither emits reasoning tokens; both answer straight away. Codestral returned code for every task, with no empty bodies and no API-layer failures, so its run is complete and 8/9 is a valid score. It is one wrong function, not a missing one.

Devstral 2 ranked 11th cheapest and 2nd fastest of the 75 models that had scored 9 out of 9 as of 2026-10-02. The only 9/9 model ahead of it on speed was Solar Mini 4, at 2 seconds. Codestral has no rank in that list because it did not clear the suite.

The miss: parse_csv_line

parse_csv_line asks for a parser that splits one CSV line into fields. The prose spec covers quoted fields, delimiters that appear inside quotes, and quotes escaped inside a quoted field, and the hidden asserts test each of those clauses. Nothing in it is algorithmically hard. It punishes a model that writes the obvious split-on-comma parser, or a half-correct quote handler, and stops reading the spec early. We have written about why this task catches models of every size in our Llama 4 Scout review.

Codestral had company on 2026-10-02:

Model (2026-10-02 sweep)ScoreMissedCost / 1k tasks (priced 2026-10-02)Latency
Codestral 25.088/9parse_csv_line$0.112.5s
Qwen3 Coder Flash8/9parse_csv_line$0.122s
Qwen3 Coder Plus8/9parse_csv_line$0.412s
GLM 5.3 Prime7/9token_bucket, parse_csv_line$11.5214.2s

Every scored model in that table missed parse_csv_line, and they range from 11 cents to $11.52 per thousand tasks. Price did not buy the fix. A fifth endpoint, the free stealth model Space Bunny Alpha, also got it wrong in one of its two runs, but both of those runs were incomplete, so it is excluded and has no valid score; it does not belong in the table beside the others, and as a free endpoint it is barred from any cost comparison anyway.

A second view from the same day. When we pushed the nine tasks through TypeSafe's Jev Router, a model-picking router, it sent parse_csv_line to Claude Opus 5.5, and that single task was 56% of the summed cost of the probe (cost as reported by the API per call, 2026-10-02). The router, too, treated this as the one task worth paying up for.

What Codestral is built for

The fair reading of an 8/9 depends on what the model was designed to do. Per Mistral's own announcement (third-party, read 2026-10-02), Codestral 25.08 was released on July 30, 2025 and is built for high-precision fill-in-the-middle (FIM) completion: the editor sends the code before and after the cursor, and the model fills the gap. Mistral's headline claims for this release are about that use case — 30% more accepted completions and 50% fewer runaway generations — with a smaller 5% instruction-following gain claimed for chat mode.

Our harness tests something different. We send a function signature and a prose spec through the ordinary chat route, the same way for every model, and ask for the whole function. There is no surrounding code to anchor on, and the spec is the only source of truth. That is a chat-mode, instruction-following job, which is the part of Codestral Mistral says improved by 5%, not the part it says improved by 30%. We measured Codestral off its home ground. It still solved eight of nine, which is a decent result for a completion model, but it is not evidence of how it performs as an autocomplete engine in an IDE.

Devstral 2 is the opposite case. TechCrunch reported on December 9, 2025 (third-party, read 2026-10-02) that Devstral 2 is a 123-billion-parameter coding model released under a modified MIT licence. It is pitched at writing and changing code from instructions — much closer to what our harness asks for.

Codestral vs Devstral 2

The surprising part of the comparison is how similar the two runs look apart from the score. Codestral wrote 950 output tokens across the suite; Devstral 2 wrote 911, 39 fewer. Neither spent a single token on reasoning. So the cost gap is almost entirely the rate card: Devstral 2 charges $2 per million output tokens against Codestral's $0.9 (both priced 2026-10-02), 2.2 times more (2 ÷ 0.9), and its measured cost per thousand tasks comes out 2.1 times higher (0.23 ÷ 0.11).

In absolute terms that multiple is 12 cents per thousand tasks ($0.23 − $0.11, priced 2026-10-02). For the extra 12 cents Devstral 2 got parse_csv_line right and answered 0.4 seconds faster at the mean. If your workload is generating functions from written requirements, that is an easy trade. Codestral being cheaper does not help much when the thing it saves you money on is a wrong answer you then have to find.

The input-token difference — 588 for Codestral, 696 for Devstral 2 on identical prompts — is 108 tokens. We did not investigate why; tokenizer or template differences between the two endpoints are plausible, but we have not confirmed either.

Where $0.11 sits

Codestral is cheap. It is not cheap for a model that misses a task.Measured cost per 1,000 tasks, priced 2026-10-02. Grey = 9/9. Rose = 8/9. Blue = Codestral 25.08 (8/9).Solar Mini 4$0.03 · 9/9Codestral 25.08$0.11 · 8/9Qwen3 Coder Flash$0.12 · 8/9Qwen3 Coder$0.13 · 9/9Devstral 2$0.23 · 9/9Qwen3 Coder Plus$0.41 · 8/9One scale throughout: 1,200 px per dollar. Bar width = cost × 1,200 ($0.03 = 36 px, $0.11 = 132 px, $0.12 = 144 px,$0.13 = 156 px, $0.23 = 276 px, $0.41 = 492 px). All six measured 2026-10-02 on the same nine tasks. Local set, not the full field.
Within these six, the two cheapest bars that are not 8/9 belong to Solar Mini 4 and Qwen3 Coder.

Two comparisons from the chart matter for anyone choosing on price. Solar Mini 4 — the cheapest and fastest 9/9 in our data as of 2026-10-02 — scored 9/9 at $0.03 and 2 seconds, so Codestral costs 3.7 times as much (0.11 ÷ 0.03) and scored lower. Qwen3 Coder scored 9/9 at $0.13 in 2.7 seconds, two cents more than Codestral (both priced 2026-10-02) for the task Codestral missed. On this suite, Codestral sits in an awkward spot: not cheap enough to be the budget pick, not complete enough to be the safe one.

Keep the scale of all this in proportion. Nine self-contained Python functions cannot separate a frontier model from a competent small one — 75 models have cleared the whole suite. A one-task difference on a single run is a real measurement, but it is a thin one. For the wider cheap end, our cheap coding roundup has the field, and our Mistral Medium 3.5 review covers Mistral's general-purpose line.

Who should use it

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. The generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why Codestral's miss counts against its score while excluded runs get no score at all. Cost is derived: measured input and output token counts multiplied by the list price captured on 2026-10-02 — it is not a billing statement, and prices move. The Jev Router figure is the exception: there the cost is what the API reported per call. All runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

What we did not measure

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.