qwen3.8-max-0902: The Dated Checkpoint Is the Slow One
qwen3.8-max-0902 scored 9 out of 9 on our executed Python suite and derived $4.21 per 1,000 tasks at the list price captured 2026-09-15. Its undated sibling qwen3.8-max scored the same 9 out of 9 and derived $4.46 at the list price captured 2026-08-20. Slightly cheaper, then, for the checkpoint with the date on it. The part nobody puts in a release note is the clock: 25.2 seconds mean against 16.8 seconds — 25.2 ÷ 16.8 = 1.5 exactly. Among the 56 models in our set that clear the whole suite, 0902 ranks 54th fastest, sitting directly ahead of qwen3.7-max at 55th, the generation it is two releases past. The suffix bought a few cents. It handed back the generation's entire latency gain.
A dated checkpoint is supposed to be the boring, reproducible option: the same weights you tested last week, pinned so a vendor rollout cannot move under your production traffic. That is a real thing to want. What is worth knowing before you pin is what the pin costs, and on our suite the answer was not in the price column.
Three generations, one harness
All three ran the identical nine executed Python tasks, all three are post-sweep entries in our data, and all three carry a 1,000,000-token context.
| Metric | Qwen3.8-Max-0902 | Qwen3.8-Max | Qwen3.7-Max |
|---|---|---|---|
| Score | 9/9 | 9/9 | 9/9 |
| Derived cost / 1,000 tasks | $4.21 · priced 2026-09-15 | $4.46 · priced 2026-08-20 | $6.1 · priced 2026-07-30 |
| Mean latency | 25.2s | 16.8s | 25.8s |
| Reasoning tokens per call | 546 | 589 | 1,236 |
| Tokens across the suite | 1,105 in / 5,951 out | 988 in / 6,363 out | 646 in / 12,193 out |
| List price when measured | $2 in / $6 out per 1M | $2 in / $6 out per 1M | $1.475 in / $4.425 out per 1M |
| Rank of the 56 at 9/9 | 37 cheapest · 54 fastest | 39 cheapest · 47 fastest | 44 cheapest · 55 fastest |
| Run date | 2026-09-16 | 2026-08-20 | 2026-07-30 |
Read the token columns before the money columns, because they drive it. Qwen3.7-Max spent 1,236 reasoning tokens per call and 12,193 output tokens across the suite. Both 3.8 checkpoints spend far less: 589 and 546 reasoning tokens per call, 6,363 and 5,951 output tokens. That is the generational change, and we wrote it up when 3.8 landed — Alibaba raised the list price 3.7 → 3.8 and the measured bill still fell, because conciseness outran the price card.
One tidy detail that survives the rewrite: both price cards hold output at exactly three times input. $6 ÷ $2 = 3, and $4.425 ÷ $1.475 = 3. Alibaba re-based the numbers to round figures without touching the shape of the card.
A warning on that 3.7-Max column. The $6.1 was derived at $1.475 in / $4.425 out, the list on 2026-07-30. Today the card lists at $1.475 in / $4.42 out — the output side moved. The distance between $4.425 and $4.42 is nothing to your invoice and everything to a reproduction attempt: re-derive with today's number and you will not land on $6.1. This is the ordinary condition of the market, not an anomaly — see how often these cards move.
The clock is where the suffix shows up
Put a number on the reference bar at the bottom: GPT-5.4-mini clears the same nine tasks in a 2.3-second mean. Both Qwen Max checkpoints sit in the bottom third of the 9/9 field for speed — 54th and 47th fastest of 56. If you are putting a Max model behind anything a human waits on, that is the constraint, not the $4.21.
We are deliberately not calling this a regression in the weights. One scored attempt per task, one run per checkpoint. A mean built that way absorbs whatever the provider's queue looked like on 2026-09-16, and 0902 also emitted the most input tokens of the three — 1,105 across the suite. What we can say is narrow and still useful: on the day we measured it, the pinned checkpoint was the slow one, and pinning did not get us the 16.8-second behaviour we recorded in August.
What the undated id does and does not tell you
The obvious question is whether qwen/qwen3.8-max simply is 0902 today, in which case our two rows are two samples of one model and the latency gap is noise. We cannot answer that from this data, and we are not going to pretend otherwise.
Here is what we actually hold. We captured a resolution map for the -latest style pointers on 2026-09-15 and published it as the alias map; every entry in it is an ~vendor/family-latest id, and there is no Qwen entry. The two Qwen rows in this article are two separate runs, on 2026-08-20 and 2026-09-16, against two separate ids. That is enough to compare what each id served us on its own date. It is not enough to claim the undated id changed, or that it did not.
Which is the whole practical case for the dated suffix, and it is a good one: the value of a pinned checkpoint is not that it is better, it is that this uncertainty stops being yours. You pay for it in whatever the frozen snapshot happens to be slower or worse at, and for 0902 the bill came in wall-clock seconds rather than dollars. If you are choosing between Qwen Max generations rather than between pins, our Qwen Max review untangles which name is which, and the 3.8 launch audit covers what was claimed versus what was verifiable at announcement.
The batch endpoint is not a safe assumption here
If your Max workload is offline, the usual move is the :batch endpoint at half price. We pulled the batch pairs on 2026-09-15: 85 base models publish one, and 70 of them price it at exactly 50% of standard output. Fifteen do not, and the deviations are not small — openai/gpt-oss-120b bills batch at 353% of its standard card, google/gemma-4-31b-it at 285%.
The Qwen 3.8 entry in that group of fifteen is qwen/qwen3.8-2.4t-a95b: standard $2.00/$6.00, batch $2.00/$6.00, a ratio of 100% — the batch endpoint exists and discounts nothing. That is the same list price the Max checkpoints carry. It is not a Max id, so we are not telling you 0902's batch endpoint behaves that way; we are telling you the 50% default is wrong often enough, and wrong within this family, that you should read the card instead of assuming. GLM-5.3-Flash sits in the same 100% bucket.
What a $4.21 model proves against a $0.07 one
Everything above is a comparison between Max checkpoints, and that is the only comparison this suite can carry honestly. Nine self-contained Python functions cannot separate a frontier model from a competent small one. Fifty-six models clear all nine, and the cheapest of them, Ling 3.0 Flash VL, does it for $0.07 per 1,000 tasks at the list price captured 2026-09-15, with a 4.4-second mean.
So the honest statement of the result is: on work of this size, 0902's 9/9 is a floor check, not a capability finding. A trillion-parameter-class Max checkpoint and a cheap flash model are indistinguishable here, which tells you the suite has a ceiling, not that the models are equivalent. If your workload really is this shape — short, self-contained, verifiable functions — then the $4.21 is buying you nothing our harness can see, and the cheap end of the field is the right place to look. If it is not this shape, our number is a sanity check and the evaluation you need is your own. The full cost benchmark has the rest of the field.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. The generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived from measured token counts at list price on the date named beside each figure — it is not a billing statement. Runs go through OpenRouter, not through our own gateway. Current list prices for these ids are on the Qwen pricing page, and the full method is on the methodology page.
What we did not measure
- Whether the undated id resolves to 0902. We ran two ids on two dates. That is not a resolution test, and our alias capture has no Qwen entry.
- Latency variance. One run per checkpoint. The 25.2s mean could be a loaded provider on 2026-09-16 rather than a property of the snapshot, and we would not know from this data.
- The 1,000,000-token context. Our entire suite fed 0902 1,105 input tokens against a 1,000,000-token window. The long-context behaviour is untested.
- The batch endpoint for either Max id. We hold the batch card for
qwen/qwen3.8-2.4t-a95b, not for these two. - Anything a Max model is bought for. Long refactors, multi-turn agent loops, tool calling, non-English work. Nine one-shot functions touch none of it.
- Any comparison against a partial run. Models whose runs failed at the API layer are marked excluded in our data and carry no score; we do not print a 7-of-7 beside a 9-of-9 as though they were the same measurement.
FAQ
What does qwen3.8-max-0902 cost per task? A derived $4.21 per 1,000 tasks, from measured token counts at the $2 in / $6 out per 1M list captured 2026-09-15.
Is the dated checkpoint better than qwen3.8-max? Not on our suite. Both scored 9 out of 9; 0902 derived $4.21 against $4.46 and posted a 25.2-second mean against 16.8 seconds.
Why pin a dated checkpoint at all? So a vendor-side rollout cannot change your production behaviour without you choosing it. That is a stability argument, not a quality one.
Is it worth upgrading from Qwen3.7-Max? On cost, yes: $6.1 derived on 2026-07-30 against $4.21 derived on 2026-09-15, driven by 12,193 output tokens falling to 5,951. On latency, 25.8s against 25.2s is not an upgrade worth the migration on its own.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab