Model Reviews

Qwen3 Coder Review: The Plus Tier Costs About 3x and Scored Lower

The priciest Qwen3 Coder tier in our run scored lower than the base model. In this Qwen3 Coder review, run on 2026-10-02 against nine executed Python tasks, the base Qwen3 Coder scored 9 out of 9 at $0.13 per 1,000 tasks with a 2.7-second mean. Qwen3 Coder Plus scored 8 out of 9 at $0.41 — about 3x the bill for one fewer correct answer — and it is not a hidden prompt or extra thinking that costs more: all three Qwen3 Coder endpoints we ran that day read the identical 619 input tokens and none emitted a single reasoning token. The difference is the rate card. The task Plus got wrong was parse_csv_line, and so did the cheaper Qwen3 Coder Flash.

DataLLM Lab article cover: Qwen3 Coder Review: The Plus Tier Costs About 3x and Scored Lower

A vendor's tier names imply an ordering: Flash below, the plain model in the middle, Plus on top. On our harness the Qwen3 Coder family does not sort that way, and the reason is narrower and more useful than “Plus is worse”.

The four models on one harness

ModelScoreCost / 1,000 tasksPriced onMean latencyTokens in / out (suite)Price used for cost, in / out per 1MContext
Qwen3 Coder9/9$0.132026-10-022.7s619 / 1,010$0.3 / $1262,144
Qwen3 Coder Plus8/9$0.412026-10-022s619 / 1,020$0.65 / $3.251,000,000
Qwen3 Coder Flash8/9$0.122026-10-022s619 / 1,020$0.195 / $0.9751,000,000
Qwen3 Coder Next9/9$0.12026-07-177snot storednot stored for that run262,144

The first three were run on the same day through the same route. Qwen3 Coder Next is a reference from our original core-13 set: its score is valid, but its cost was derived at a price captured 2026-07-17 and its latency comes from a different day, so treat that row as context, not as a same-day comparison.

Two results in that table are complete runs with a genuine wrong answer — every task returned code, the code executed, and one assert failed. That makes 8/9 a valid score, not an exclusion. Among the 75 models that scored 9 out of 9 as of 2026-10-02, the base Qwen3 Coder ranks 8th cheapest and 6th fastest.

Plus: 3x the bill, one task fewer

We have written before about a premium tier that costs more because it silently prepends a large prompt to every request. That is not what happens here. Qwen3 Coder, Plus and Flash each read 619 input tokens across the nine prompts — identical, so no tier is adding scaffolding — and each reported 0 reasoning tokens per call. Output was 1,010 tokens for the base model and 1,020 for Plus: ten tokens apart over nine tasks.

So the whole gap is the rate card. Plus lists at $0.65 in and $3.25 out per million as of 2026-10-02; the base model at $0.3 and $1. On a workload that is mostly output, the output rate decides the bill, and $3.25 ÷ $1 is 3.25x. Measured: $0.41 ÷ $0.13 = 3.2x per 1,000 tasks, both priced 2026-10-02.

Be precise about what the score difference means. It is one task, on one scored attempt, at temperature 0. A re-run could plausibly flip it. What we would not expect a re-run to change is the price: with input and output token counts this close, Plus will cost roughly three times the base model on short coding prompts regardless of who wins the ninth task. The score is the noisy half of this result. The cost multiple is the stable half.

parse_csv_line is where the family splits

Both Qwen3 Coder Plus and Qwen3 Coder Flash missed the same task, parse_csv_line, and posted identical usage counts: 619 tokens in, 1,020 out, a 2s mean. That is an observation about two endpoints, not a claim that they serve the same weights — we did not diff the generated code, and we have been wrong before inferring identity from behaviour. It is consistent with OpenRouter's description of Flash as a cheaper version of Plus (third-party, read 2026-10-02).

They were not alone. On the same day:

ModelScoreMissedCost / 1,000 tasks (priced 2026-10-02)
Qwen3 Coder Plus8/9parse_csv_line$0.41
Qwen3 Coder Flash8/9parse_csv_line$0.12
Codestral 25088/9parse_csv_line$0.11
GLM-5.3 Prime7/9token_bucket, parse_csv_line$11.52
Command A PlusExcluded — no valid score. Its parse_csv_line call returned an empty body with finish_reason=length, an API-layer failure rather than a wrong answer, so only 8 of 9 tasks were scored and the run cannot be ranked.

The task is a single-line CSV parser, and its spec is the kind where quoting rules matter more than algorithmic skill. A separate probe the same day points the same way: a per-request router running our nine tasks sent parse_csv_line to Claude Opus 5.5, and that one call was 56% of the nine-task bill. The task that splits the Qwen3 Coder family is the one where that router spent the most of its nine calls.

Where the family sits on cost

The top Qwen3 Coder tier has the longest bar and a lower scoreMeasured cost per 1,000 tasks. Blue: Qwen3 Coder at 9/9. Amber: Qwen3 Coder at 8/9. Grey: same-day references.Solar Mini 4 · 9/9$0.03Qwen3 Coder Next · 9/9$0.1 · priced 2026-07-17Qwen3 Coder Flash · 8/9$0.12Qwen3 Coder · 9/9$0.13Devstral 2512 · 9/9$0.23Qwen3 Coder Plus · 8/9$0.41Scale: 1,250 px per dollar; bar width = cost × 1,250 ($0.41 = 512.5 px). Priced 2026-10-02 except Next.
Every bar uses the same 1,250 px per dollar. The family's two 8/9 tiers sit at both ends of its price range.

For orientation: Solar Mini 4 is the cheapest and fastest model to score 9 out of 9 as of 2026-10-02, at $0.03 and a 2s mean. Devstral 2512 scored 9/9 the same day at $0.23 and 2.1s. The base Qwen3 Coder sits between them on cost and is the one Qwen3 Coder that both cleared the suite and was measured on the current date.

Flash is the awkward middle. At $0.12 it costs one cent less than the base model, priced 2026-10-02, and lost a task. At the 2026-10-02 run prices our costs use, Flash was exactly 30% of Plus on both sides — $0.195 against $0.65 in, $0.975 against $3.25 out — and since the two produced identical token counts, Plus cost $0.41 ÷ $0.12 = 3.4x Flash on the rounded figures for the same 8 out of 9. Flash's output list price has since moved to $0.97, so do not apply that ratio to today's rate card.

Qwen3 Coder Next: cheaper, but an older measurement

Qwen3 Coder Next scored 9 out of 9 at $0.1 per 1,000 tasks, priced 2026-07-17, which ranks it 5th cheapest of the 75 models at 9/9 as of 2026-10-02. It ran at a 7s mean — 37th fastest — against 2.7s for the base Qwen3 Coder on 2026-10-02.

We would not read too much into either gap. That run predates the field where we store the price a cost was computed from, and the token counts were not kept, so we cannot re-derive it at today's rate. For the record, the catalogue lists Next at $0.12 in and $0.8 out per million as of 2026-10-02; we have not re-priced the old run at that rate, and the $0.1 above should be read with its 2026-07-17 date attached. If Next is your candidate, it deserves a same-day re-run before you choose on three cents.

What the family is, according to third parties

None of the following is our measurement. From OpenRouter's model catalogue and the Qwen repositories on Hugging Face, both read 2026-10-02:

The practical consequence is that only the base model and Next can be self-hosted. If you want to run the open base model yourself, our Ollama guide for Qwen3 Coder covers the tags and the memory it needs; Plus and Flash are API-only.

Which one to use

For the cheap end beyond Qwen, see our cheap coding roundup; for the family's other price lists, Qwen API pricing.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why Plus and Flash carry a valid 8/9 while Command A Plus is excluded. Cost is derived from measured token counts at the list price captured on the date shown, not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

Be clear about the ceiling of this instrument. Nine self-contained Python functions cannot separate a frontier coding model from a competent small one; 75 of the 93 usable models in our set score 9 out of 9 as of 2026-10-02. What the suite does well is expose price per correct function and the occasional task a model reliably fumbles.

What we did not measure

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.