Model Reviews

Qwen3.8 Max Prime Review: Twice the Price, 1.3 Seconds Faster

Qwen3.8 Max Prime is sold as the fast lane for Qwen3.8 Max. On our executed Python benchmark it delivered a 15.5-second mean against 16.8 seconds for plain qwen3.8-max — 16.8 − 15.5 = 1.3 seconds. Both scored 9 out of 9. Prime derived $12.41 per 1,000 tasks at the list price captured 2026-10-02; qwen3.8-max derived $4.46 at the list price captured 2026-08-20. That is 12.41 ÷ 4.46 = 2.78x the bill on a rate card that is only 2x. The missing multiple is behaviour, not pricing: on the same nine prompts Prime wrote 8,938 output tokens where qwen3.8-max wrote 6,363.

DataLLM Lab article cover: Qwen3.8 Max Prime Review: Twice the Price, 1.3 Seconds Faster

A speed tier is a simple promise: same model, same answers, delivered sooner, for more money. We measured all three halves of that promise. The answers were the same. The delivery was barely sooner. The money was not simple at all.

The result

MetricQwen3.8 Max Primeqwen3.8-maxqwen3.8-max-0902
Score9/99/99/9
Measured cost / 1,000 tasks$12.41$4.46$4.21
List price priced on2026-10-022026-08-202026-09-15
Priced at, in / out per 1M$4 / $12$2 / $6$2 / $6
Mean latency15.5s16.8s25.2s
Reasoning tokens per call875589546
Input tokens across the suite1,1059881,105
Output tokens across the suite8,9386,3635,951
Context window1,000,0001,000,0001,000,000
Rank of 75 at 9/9: cheapest / fastest72 / 6252 / 6550 / 72
Run date2026-10-022026-08-202026-09-16

Both older list prices were unchanged as of 2026-10-02, so the cost column compares like with like even though the three runs happened on different days. Among the 75 models in our set that clear the full suite, Prime ranks 72nd cheapest. It is also the most expensive of the six Qwen endpoints in the family table further down — that superlative is scoped to those six, not to everything Alibaba has ever shipped.

What the speed tier bought

Third-party context first, labelled as such. According to OrcaRouter's write-up (read 2026-10-02), which cites Alibaba's own documentation, Prime was announced at Alibaba's Yunqi Conference in Hangzhou in September as a serving tier: the same weights, context window and tooling as Qwen3.8 Max, with output throughput of 1.5 to 2 times the standard API. models.dev (read 2026-10-02) lists the weights as closed and the context at 1,000,000 tokens, which matches what the endpoint reported to us. We have not verified any of those claims beyond what our own runs show.

What our runs show is a mean that moved from 16.8 to 15.5 seconds, and a rank that moved from 65th fastest to 62nd of 75. Against the dated checkpoint qwen3.8-max-0902 the gap is larger — 25.2 − 15.5 = 9.7 seconds — but that comparison is the one we trust least, for reasons in the limits section.

A throughput claim and a latency measurement are not the same thing. If Prime really streams tokens 1.5 to 2 times faster, but also writes more of them before it stops, the wall clock can barely move. That is consistent with what we see. It is not proof of it.

The speed tier moved the clock 1.3 seconds and the bill 2.78x.Five models, all 9/9 on the same nine executed Python tasks.MEASURED COST PER 1,000 TASKSSolar Mini 4$0.03Qwen3 Coder$0.13qwen3.8-max-0902$4.21qwen3.8-max$4.46Qwen3.8 Max Prime$12.41MEAN LATENCY PER CALLSolar Mini 42sQwen3 Coder2.7sqwen3.8-max-090225.2sqwen3.8-max16.8sQwen3.8 Max Prime15.5sScale: cost 45 px per $1, so $12.41 = 558.45 px; latency 20 px per second, so 25.2s = 504 px.Costs derived from measured tokens at list price on each priced-on date: 2026-10-02 for Prime, Qwen3 Coder and Solar Mini 4;2026-08-20 for qwen3.8-max; 2026-09-15 for qwen3.8-max-0902. Not a billing statement.
Prime is the longest cost bar and only the third-longest latency bar. Solar Mini 4 and Qwen3 Coder are drawn to the same scale.

Why the bill is 2.78x, not 2x

The rate card explains exactly half of it. Prime is priced at $4 in and $12 out per million; qwen3.8-max at $2 and $6. 4 ÷ 2 = 2 and 12 ÷ 6 = 2 — a clean doubling on both sides, as of 2026-10-02.

The rest is tokens. On the identical nine prompts, at temperature 0:

Twice the price, multiplied by roughly 1.40x the output, is why the measured bill lands at 2.78x rather than 2x. That is the pattern we keep running into, and wrote up in our agent-cost piece: the rate card sets the unit price, the model sets the quantity, and the quantity moves more than anyone budgets for.

One oddity we can report but not explain: Prime's input count, 1,105, matches qwen3.8-max-0902 exactly and not the undated qwen3.8-max at 988. The prompts were byte-identical, so the difference sits in how each endpoint wraps them. We do not know what that wrapping is, and we are not going to guess which checkpoint Prime serves from a token count.

Against the rest of the Qwen line-up

EndpointScoreCost / 1k tasksPriced onPriced at in / outMean latency
Qwen3.8 Max Prime9/9$12.412026-10-02$4 / $1215.5s
qwen3.8-max9/9$4.462026-08-20$2 / $616.8s
qwen3.8-max-09029/9$4.212026-09-15$2 / $625.2s
Qwen3 Coder Plus8/9$0.412026-10-02$0.65 / $3.252s
Qwen3 Coder9/9$0.132026-10-02$0.3 / $12.7s
Qwen3 Coder Flash8/9$0.122026-10-02$0.195 / $0.9752s

Both 8/9 Coder variants missed the same task, parse_csv_line. Qwen3 Coder Flash's list price has moved since its run; the cost shown is computed from the price on the run date, 2026-10-02.

The plain Qwen3 Coder scored the same 9 out of 9 as Prime at $0.13, priced 2026-10-02, in a 2.7-second mean and with zero reasoning tokens. 12.41 ÷ 0.13 = 95x. Across the whole field, the cheapest and fastest 9/9 is Solar Mini 4 at $0.03 and 2 seconds, priced 2026-10-02.

Is Prime worth $12.41?

On this suite, no, and it would be strange if it were. Our nine self-contained Python functions cannot separate a frontier model from a competent small one — Qwen3 Coder proves it from inside the same vendor. A 9/9 here tells you a model clears the bar, not how far over it.

The narrower question is whether Prime beats its own base tier, which is the only comparison the product is selling. Our answer is: identical score, 1.3 seconds of mean latency, 2.78x the derived cost. If your workload is a short, single-turn request, the time saved is smaller than the network jitter you already live with. If your workload is long streamed generations, the throughput claim may matter far more than our mean suggests — but that is a workload you should measure, not one we did.

If you run Qwen3.8 Max today and latency is not your binding constraint, the base qwen3.8-max is the same answer at $4.46 per 1,000 tasks, priced 2026-08-20, against $12.41, priced 2026-10-02. Our Qwen pricing guide covers the other tiers.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; Prime had all nine tasks scored. Cost is derived — measured input and output token counts multiplied by the list price on the date shown, not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

What we did not measure

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.