Qwen3 Coder Review: The Plus Tier Costs About 3x and Scored Lower
The priciest Qwen3 Coder tier in our run scored lower than the base model. In this Qwen3 Coder review, run on 2026-10-02 against nine executed Python tasks, the base Qwen3 Coder scored 9 out of 9 at $0.13 per 1,000 tasks with a 2.7-second mean. Qwen3 Coder Plus scored 8 out of 9 at $0.41 — about 3x the bill for one fewer correct answer — and it is not a hidden prompt or extra thinking that costs more: all three Qwen3 Coder endpoints we ran that day read the identical 619 input tokens and none emitted a single reasoning token. The difference is the rate card. The task Plus got wrong was parse_csv_line, and so did the cheaper Qwen3 Coder Flash.
A vendor's tier names imply an ordering: Flash below, the plain model in the middle, Plus on top. On our harness the Qwen3 Coder family does not sort that way, and the reason is narrower and more useful than “Plus is worse”.
The four models on one harness
| Model | Score | Cost / 1,000 tasks | Priced on | Mean latency | Tokens in / out (suite) | Price used for cost, in / out per 1M | Context |
|---|---|---|---|---|---|---|---|
| Qwen3 Coder | 9/9 | $0.13 | 2026-10-02 | 2.7s | 619 / 1,010 | $0.3 / $1 | 262,144 |
| Qwen3 Coder Plus | 8/9 | $0.41 | 2026-10-02 | 2s | 619 / 1,020 | $0.65 / $3.25 | 1,000,000 |
| Qwen3 Coder Flash | 8/9 | $0.12 | 2026-10-02 | 2s | 619 / 1,020 | $0.195 / $0.975 | 1,000,000 |
| Qwen3 Coder Next | 9/9 | $0.1 | 2026-07-17 | 7s | not stored | not stored for that run | 262,144 |
The first three were run on the same day through the same route. Qwen3 Coder Next is a reference from our original core-13 set: its score is valid, but its cost was derived at a price captured 2026-07-17 and its latency comes from a different day, so treat that row as context, not as a same-day comparison.
Two results in that table are complete runs with a genuine wrong answer — every task returned code, the code executed, and one assert failed. That makes 8/9 a valid score, not an exclusion. Among the 75 models that scored 9 out of 9 as of 2026-10-02, the base Qwen3 Coder ranks 8th cheapest and 6th fastest.
Plus: 3x the bill, one task fewer
We have written before about a premium tier that costs more because it silently prepends a large prompt to every request. That is not what happens here. Qwen3 Coder, Plus and Flash each read 619 input tokens across the nine prompts — identical, so no tier is adding scaffolding — and each reported 0 reasoning tokens per call. Output was 1,010 tokens for the base model and 1,020 for Plus: ten tokens apart over nine tasks.
So the whole gap is the rate card. Plus lists at $0.65 in and $3.25 out per million as of 2026-10-02; the base model at $0.3 and $1. On a workload that is mostly output, the output rate decides the bill, and $3.25 ÷ $1 is 3.25x. Measured: $0.41 ÷ $0.13 = 3.2x per 1,000 tasks, both priced 2026-10-02.
Be precise about what the score difference means. It is one task, on one scored attempt, at temperature 0. A re-run could plausibly flip it. What we would not expect a re-run to change is the price: with input and output token counts this close, Plus will cost roughly three times the base model on short coding prompts regardless of who wins the ninth task. The score is the noisy half of this result. The cost multiple is the stable half.
parse_csv_line is where the family splits
Both Qwen3 Coder Plus and Qwen3 Coder Flash missed the same task, parse_csv_line, and posted identical usage counts: 619 tokens in, 1,020 out, a 2s mean. That is an observation about two endpoints, not a claim that they serve the same weights — we did not diff the generated code, and we have been wrong before inferring identity from behaviour. It is consistent with OpenRouter's description of Flash as a cheaper version of Plus (third-party, read 2026-10-02).
They were not alone. On the same day:
| Model | Score | Missed | Cost / 1,000 tasks (priced 2026-10-02) |
|---|---|---|---|
| Qwen3 Coder Plus | 8/9 | parse_csv_line | $0.41 |
| Qwen3 Coder Flash | 8/9 | parse_csv_line | $0.12 |
| Codestral 2508 | 8/9 | parse_csv_line | $0.11 |
| GLM-5.3 Prime | 7/9 | token_bucket, parse_csv_line | $11.52 |
| Command A Plus | Excluded — no valid score. Its parse_csv_line call returned an empty body with finish_reason=length, an API-layer failure rather than a wrong answer, so only 8 of 9 tasks were scored and the run cannot be ranked. | ||
The task is a single-line CSV parser, and its spec is the kind where quoting rules matter more than algorithmic skill. A separate probe the same day points the same way: a per-request router running our nine tasks sent parse_csv_line to Claude Opus 5.5, and that one call was 56% of the nine-task bill. The task that splits the Qwen3 Coder family is the one where that router spent the most of its nine calls.
Where the family sits on cost
For orientation: Solar Mini 4 is the cheapest and fastest model to score 9 out of 9 as of 2026-10-02, at $0.03 and a 2s mean. Devstral 2512 scored 9/9 the same day at $0.23 and 2.1s. The base Qwen3 Coder sits between them on cost and is the one Qwen3 Coder that both cleared the suite and was measured on the current date.
Flash is the awkward middle. At $0.12 it costs one cent less than the base model, priced 2026-10-02, and lost a task. At the 2026-10-02 run prices our costs use, Flash was exactly 30% of Plus on both sides — $0.195 against $0.65 in, $0.975 against $3.25 out — and since the two produced identical token counts, Plus cost $0.41 ÷ $0.12 = 3.4x Flash on the rounded figures for the same 8 out of 9. Flash's output list price has since moved to $0.97, so do not apply that ratio to today's rate card.
Qwen3 Coder Next: cheaper, but an older measurement
Qwen3 Coder Next scored 9 out of 9 at $0.1 per 1,000 tasks, priced 2026-07-17, which ranks it 5th cheapest of the 75 models at 9/9 as of 2026-10-02. It ran at a 7s mean — 37th fastest — against 2.7s for the base Qwen3 Coder on 2026-10-02.
We would not read too much into either gap. That run predates the field where we store the price a cost was computed from, and the token counts were not kept, so we cannot re-derive it at today's rate. For the record, the catalogue lists Next at $0.12 in and $0.8 out per million as of 2026-10-02; we have not re-priced the old run at that rate, and the $0.1 above should be read with its 2026-07-17 date attached. If Next is your candidate, it deserves a same-day re-run before you choose on three cents.
What the family is, according to third parties
None of the following is our measurement. From OpenRouter's model catalogue and the Qwen repositories on Hugging Face, both read 2026-10-02:
- Qwen3 Coder on OpenRouter is the open-weight Qwen3 Coder instruct model, a mixture-of-experts design whose weights are published on Hugging Face.
- Qwen3 Coder Plus is described by OpenRouter as Alibaba's proprietary version of that open model. No open weights.
- Qwen3 Coder Flash is described as a fast, cost-efficient version of the proprietary Plus.
- Qwen3 Coder Next is open-weight, a sparse mixture-of-experts model.
The practical consequence is that only the base model and Next can be self-hosted. If you want to run the open base model yourself, our Ollama guide for Qwen3 Coder covers the tags and the memory it needs; Plus and Flash are API-only.
Which one to use
- Default to the base Qwen3 Coder for short, self-contained code generation. It cleared the suite at $0.13, priced 2026-10-02, in 2.7s, and its weights are open.
- Pay for Plus only for something our suite cannot see, such as the 1,000,000-token context against 262,144 for the base model. On nine short functions it bought nothing measurable and cost 3.2x.
- Skip Flash on this evidence. One cent cheaper than the base model and one task worse is not a trade we can recommend from a single run.
- Re-test Next before trusting its July cost against the others' October figures.
For the cheap end beyond Qwen, see our cheap coding roundup; for the family's other price lists, Qwen API pricing.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why Plus and Flash carry a valid 8/9 while Command A Plus is excluded. Cost is derived from measured token counts at the list price captured on the date shown, not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
Be clear about the ceiling of this instrument. Nine self-contained Python functions cannot separate a frontier coding model from a competent small one; 75 of the 93 usable models in our set score 9 out of 9 as of 2026-10-02. What the suite does well is expose price per correct function and the occasional task a model reliably fumbles.
What we did not measure
- Why Plus failed
parse_csv_line. We store the verdict, not a diagnosis. We did not read the failing function, so we cannot tell you whether it mishandled quoting, escaping or whitespace. - Whether the miss is stable. One attempt at temperature 0. We did not re-run Plus or Flash to see if the failure repeats.
- Agentic coding, tool calling and multi-file edits — the work all four are marketed for. Our tasks are single-turn.
- The 1,000,000-token context of Plus and Flash, which is the most concrete thing their price buys. Our prompts are short.
- A same-day run of Next. Its row is from an earlier sweep, with no stored token counts.
- Provider variance. We measured whichever OpenRouter provider served each call; another host could differ in latency and, conceivably, output.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab