Qwen3.8 Max Prime Review: Twice the Price, 1.3 Seconds Faster
Qwen3.8 Max Prime is sold as the fast lane for Qwen3.8 Max. On our executed Python benchmark it delivered a 15.5-second mean against 16.8 seconds for plain qwen3.8-max — 16.8 − 15.5 = 1.3 seconds. Both scored 9 out of 9. Prime derived $12.41 per 1,000 tasks at the list price captured 2026-10-02; qwen3.8-max derived $4.46 at the list price captured 2026-08-20. That is 12.41 ÷ 4.46 = 2.78x the bill on a rate card that is only 2x. The missing multiple is behaviour, not pricing: on the same nine prompts Prime wrote 8,938 output tokens where qwen3.8-max wrote 6,363.
A speed tier is a simple promise: same model, same answers, delivered sooner, for more money. We measured all three halves of that promise. The answers were the same. The delivery was barely sooner. The money was not simple at all.
The result
| Metric | Qwen3.8 Max Prime | qwen3.8-max | qwen3.8-max-0902 |
|---|---|---|---|
| Score | 9/9 | 9/9 | 9/9 |
| Measured cost / 1,000 tasks | $12.41 | $4.46 | $4.21 |
| List price priced on | 2026-10-02 | 2026-08-20 | 2026-09-15 |
| Priced at, in / out per 1M | $4 / $12 | $2 / $6 | $2 / $6 |
| Mean latency | 15.5s | 16.8s | 25.2s |
| Reasoning tokens per call | 875 | 589 | 546 |
| Input tokens across the suite | 1,105 | 988 | 1,105 |
| Output tokens across the suite | 8,938 | 6,363 | 5,951 |
| Context window | 1,000,000 | 1,000,000 | 1,000,000 |
| Rank of 75 at 9/9: cheapest / fastest | 72 / 62 | 52 / 65 | 50 / 72 |
| Run date | 2026-10-02 | 2026-08-20 | 2026-09-16 |
Both older list prices were unchanged as of 2026-10-02, so the cost column compares like with like even though the three runs happened on different days. Among the 75 models in our set that clear the full suite, Prime ranks 72nd cheapest. It is also the most expensive of the six Qwen endpoints in the family table further down — that superlative is scoped to those six, not to everything Alibaba has ever shipped.
What the speed tier bought
Third-party context first, labelled as such. According to OrcaRouter's write-up (read 2026-10-02), which cites Alibaba's own documentation, Prime was announced at Alibaba's Yunqi Conference in Hangzhou in September as a serving tier: the same weights, context window and tooling as Qwen3.8 Max, with output throughput of 1.5 to 2 times the standard API. models.dev (read 2026-10-02) lists the weights as closed and the context at 1,000,000 tokens, which matches what the endpoint reported to us. We have not verified any of those claims beyond what our own runs show.
What our runs show is a mean that moved from 16.8 to 15.5 seconds, and a rank that moved from 65th fastest to 62nd of 75. Against the dated checkpoint qwen3.8-max-0902 the gap is larger — 25.2 − 15.5 = 9.7 seconds — but that comparison is the one we trust least, for reasons in the limits section.
A throughput claim and a latency measurement are not the same thing. If Prime really streams tokens 1.5 to 2 times faster, but also writes more of them before it stops, the wall clock can barely move. That is consistent with what we see. It is not proof of it.
Why the bill is 2.78x, not 2x
The rate card explains exactly half of it. Prime is priced at $4 in and $12 out per million; qwen3.8-max at $2 and $6. 4 ÷ 2 = 2 and 12 ÷ 6 = 2 — a clean doubling on both sides, as of 2026-10-02.
The rest is tokens. On the identical nine prompts, at temperature 0:
- Output: 8,938 against 6,363 — 8,938 ÷ 6,363 = 1.40x. Against the 0902 checkpoint, 8,938 ÷ 5,951 = 1.50x.
- Reasoning: 875 tokens per call against 589 — 875 ÷ 589 = 1.49x. Prime thought harder about problems the cheaper tier already solved.
- Input: 1,105 against 988. Output tokens cost three times as much as input here, so the input gap barely registers.
Twice the price, multiplied by roughly 1.40x the output, is why the measured bill lands at 2.78x rather than 2x. That is the pattern we keep running into, and wrote up in our agent-cost piece: the rate card sets the unit price, the model sets the quantity, and the quantity moves more than anyone budgets for.
One oddity we can report but not explain: Prime's input count, 1,105, matches qwen3.8-max-0902 exactly and not the undated qwen3.8-max at 988. The prompts were byte-identical, so the difference sits in how each endpoint wraps them. We do not know what that wrapping is, and we are not going to guess which checkpoint Prime serves from a token count.
Against the rest of the Qwen line-up
| Endpoint | Score | Cost / 1k tasks | Priced on | Priced at in / out | Mean latency |
|---|---|---|---|---|---|
| Qwen3.8 Max Prime | 9/9 | $12.41 | 2026-10-02 | $4 / $12 | 15.5s |
| qwen3.8-max | 9/9 | $4.46 | 2026-08-20 | $2 / $6 | 16.8s |
| qwen3.8-max-0902 | 9/9 | $4.21 | 2026-09-15 | $2 / $6 | 25.2s |
| Qwen3 Coder Plus | 8/9 | $0.41 | 2026-10-02 | $0.65 / $3.25 | 2s |
| Qwen3 Coder | 9/9 | $0.13 | 2026-10-02 | $0.3 / $1 | 2.7s |
| Qwen3 Coder Flash | 8/9 | $0.12 | 2026-10-02 | $0.195 / $0.975 | 2s |
Both 8/9 Coder variants missed the same task, parse_csv_line. Qwen3 Coder Flash's list price has moved since its run; the cost shown is computed from the price on the run date, 2026-10-02.
The plain Qwen3 Coder scored the same 9 out of 9 as Prime at $0.13, priced 2026-10-02, in a 2.7-second mean and with zero reasoning tokens. 12.41 ÷ 0.13 = 95x. Across the whole field, the cheapest and fastest 9/9 is Solar Mini 4 at $0.03 and 2 seconds, priced 2026-10-02.
Is Prime worth $12.41?
On this suite, no, and it would be strange if it were. Our nine self-contained Python functions cannot separate a frontier model from a competent small one — Qwen3 Coder proves it from inside the same vendor. A 9/9 here tells you a model clears the bar, not how far over it.
The narrower question is whether Prime beats its own base tier, which is the only comparison the product is selling. Our answer is: identical score, 1.3 seconds of mean latency, 2.78x the derived cost. If your workload is a short, single-turn request, the time saved is smaller than the network jitter you already live with. If your workload is long streamed generations, the throughput claim may matter far more than our mean suggests — but that is a workload you should measure, not one we did.
If you run Qwen3.8 Max today and latency is not your binding constraint, the base qwen3.8-max is the same answer at $4.46 per 1,000 tasks, priced 2026-08-20, against $12.41, priced 2026-10-02. Our Qwen pricing guide covers the other tiers.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; Prime had all nine tasks scored. Cost is derived — measured input and output token counts multiplied by the list price on the date shown, not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
What we did not measure
- Throughput, which is the product. We record mean wall-clock latency per call and token totals per suite. Dividing one by the other would produce a number that looks like tokens per second and is not one, so we have not printed it. Prime's 1.5-to-2x claim remains Alibaba's, unverified by us.
- Same-day latency. Prime ran on 2026-10-02, qwen3.8-max on 2026-08-20, the 0902 checkpoint on 2026-09-16. Provider load differs by day. A 1.3-second gap measured weeks apart is weak evidence in either direction; a fair test runs both tiers back to back.
- Variance. One scored attempt per task. We cannot say whether 875 reasoning tokens per call is Prime's habit or one run's draw.
- Why Prime reasons more. If it truly serves the same weights, a different default reasoning effort on the endpoint is one plausible cause. We have not tested that, and we have been wrong about endpoint behaviour before.
- Time to first token, long outputs, the 1,000,000-token context, and tool calling. Our prompts are short and our answers are single functions.
- Alibaba's direct pricing. We priced the OpenRouter list. Buying from Alibaba Cloud may cost something different.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab