Model Review

Qwen 3.8 Max Preview (2.4T): a frontier claim shipped without a single benchmark

On Sunday 19 July 2026, at the World AI Conference in Shanghai, Alibaba's Qwen team previewed Qwen3.8-Max-Preview and said it is second only to Claude Fable 5. It shipped that sentence with no model card, no benchmark table, no license, no activated-parameter count and no first-party per-token price. Most write-ups have reprinted the claim and stopped at no benchmarks yet. This one measures the claim instead: we pulled the predecessor's independent score and worked out exactly how large a jump Alibaba is silently asserting.

Qwen 3.8 Max Preview announcement audited against independently verifiable evidence, July 2026

What Alibaba actually announced

On Sunday 19 July 2026, during the World AI Conference in Shanghai, Alibaba's Qwen team previewed Qwen3.8-Max-Preview. The announcement was a social post, published in identical wording from the @Alibaba_Qwen and @alibaba_cloud accounts, reported by SCMP, the-decoder and MarkTechPost the same day.

The post says the model is launching and going open-weight soon, gives a figure of 2.4 trillion parameters, says the model is continuously evolving, and claims it is comparable to leading frontier models and second only to Fable 5. (The original writes compatible where comparable is clearly meant.) That is the entire technical disclosure. There is no attached model card, no benchmark table, no evaluation methodology, no license file and no repository.

Here is what is on the record, with its provenance labelled:

And the part that matters most: as of 21 July 2026, Qwen3.8-Max does not appear anywhere on the Artificial Analysis Intelligence Index leaderboard. We fetched it directly. The board lists 186 models. Qwen entries stop at Qwen3.7 Max, Qwen3.6 Plus and the Qwen3.5 variants. No independent evaluation of Qwen3.8 exists yet, from anyone.

Announced vs evidenced: the audit

The most useful thing we can do with an announcement like this is line it up attribute by attribute against what a buyer can independently verify. One row per attribute, one column for the claim, one column for the evidence, one column for where you can check it yourself.

Attribute What Alibaba announced (19 Jul 2026) Independent evidence (21 Jul 2026) Where to check
Total parameters 2.4 trillion None No model card has been published
Activated parameters Not disclosed None Serving cost and speed are uncomputable without it
Architecture Not stated; MoE inferred by reporters None Announcement text only
Modalities Text, image, video, documents None No public API to test outside preview
Context window Not stated on any Alibaba page ~983,616 tokens via integration metadata only help.aliyun.com and docs.qwencloud.com omit it
Quality rank Second only to Claude Fable 5 Absent from all 186 models on the Artificial Analysis leaderboard artificialanalysis.ai/leaderboards/models
Price Credits at as low as 10% of the standard rate No standard rate ever published by Alibaba; reseller rates disagree ~9x Discount is off an unknown base
Weights Open-weight soon No date, no license, no Hugging Face repo Qwen Hugging Face org
Availability Preview only Not on the international Token Plan model list alibabacloud.com Token Plan model table stops at qwen3.7-max

Nine attributes. Nine cells in the evidence column that are empty, inferred or contradicted. That is not a normal condition for a model being publicly ranked second in the world, and it is why we are not writing a review — there is nothing to review. If you want a model page where the numbers actually exist, our Qwen3-Max review covers the predecessor line that has real published scores and a real rate card.

How big the claim actually is

Here is the number nobody else pulled. Qwen3.7-Max — the direct predecessor — scores 46 on the Artificial Analysis Intelligence Index, ranked #27 of 186, priced at $2.50 per million input and $7.50 per million output with a 1M context. We fetched that page directly on 21 July 2026.

Now put that next to the top of the same board on the same day: Claude Fable 5 (with fallback) 60, GPT-5.6 Sol (max) 59, GPT-5.6 Sol (xhigh) 58, Kimi K3 57, GPT-5.6 Sol (high) 56, Claude Opus 4.8 (max) 56.

For second only to Fable 5 to be true on this index, Qwen3.8-Max has to score above Kimi K3's 57 and below Fable 5's 60. Strictly, to be second it also has to clear GPT-5.6 Sol at 59. That is a single-generation jump of roughly +12 to +13 points, and a move from rank #27 to rank #2.

That is the information gain. Every other outlet reprinted the sentence. The sentence is not vague — it is a precise, arithmetic, falsifiable assertion about a public leaderboard, and it is a very large one.

What the Qwen3.8 claim would have to be worth Artificial Analysis Intelligence Index, leaderboard read directly on 2026-07-21. Qwen3.8-Max has no score. Claude Fable 5 · #1 GPT-5.6 Sol max · #2 Kimi K3 · #4 Qwen3.8-Max · unranked Qwen3.7-Max · #27 60 59 57 46 ? not measured 40 45 50 55 60 Axis truncated at 40. The dashed bar is arithmetic implied by the vendor claim, not a measured result. Chart: DataLLM Lab
Measured points are independent (Artificial Analysis, fetched 2026-07-21). The dashed band is illustrative — it shows the interval the claim implies, not a score anyone has recorded. Chart: DataLLM Lab

Two things are worth saying plainly about that dashed band. First, it is not a prediction — we are not asserting Qwen3.8 will or will not land there. Second, a +12-point generational jump is not impossible; it is just extraordinary, and extraordinary jumps are exactly the ones that get published with evaluation harnesses attached. This one wasn't.

The pricing problem: Credits, not tokens

You will read that Qwen3.8-Max-Preview has no pricing. That is slightly too strong, and the accurate version is more interesting.

Alibaba has published no pay-as-you-go per-million-token rate card for it, and the model is absent from Alibaba Cloud's international Token Plan model table (which stops at qwen3.7-max, qwen3.7-plus, qwen3.6-plus and qwen3.6-flash, alongside DeepSeek, Moonshot, Zhipu and MiniMax models). But it is listed on the Model Studio / Qwen Cloud Personal Token Plan, with real subscription tiers:

Note the unit: Credits, not tokens. The documentation adds that during the promotional period, Credits consumption is as low as 10% of the standard rate, effectively providing 10x the usage, with a further off-peak discount for calls between 22:00 and 08:00 UTC+8. (The page's compounded arithmetic does not reconcile, so we are not quoting a combined percentage.) Access also runs through the Qoder and QoderWork agentic platforms.

Third-party gateways have already begun listing per-token rates for qwen3.8-max-preview, and they do not agree with each other. NanoGPT lists $1.50 per million input and $5.00 per million output (plus $0.15 cache read); AIHubMix lists $0.17 and $0.51, roughly a tenth of NanoGPT's figures — consistent with the 10% preview discount being baked into one listing and not the other, though neither vendor says so outright. That is a roughly 9x spread between two resellers of the same preview model, with no first-party rate card to reconcile them against. The disagreement is the finding: nobody outside Alibaba knows the standard rate, so nobody can tell you what the discount is a discount from.

So the honest framing for anyone doing capacity planning: access is metered in opaque subscription Credits, at a temporary discount off a standard rate Alibaba has never published for this model, and the only per-token numbers in circulation are reseller listings that differ by an order of magnitude. You cannot derive a defensible cost per million tokens. You therefore cannot derive a cost per task, cannot compare it to anything in your stack, and cannot put a line item in a budget. For an engineering team, that is a harder blocker than the missing benchmarks — a quality claim you can eventually test, but a unit price that does not exist cannot be modelled at all. For comparison, our Qwen API pricing breakdown shows what the published Qwen tiers actually cost per token.

For context on the international side, Alibaba's Team Edition Token Plan runs USD 30/seat/month for 25,000 Credits (Standard), USD 100 for 100,000 (Pro) and USD 200 for 250,000 (Max) — but qwen3.8-max-preview is not on that plan's model list.

One aside worth noticing: search results for this model are already flooded with pages titled things like Qwen 3.8 Max REVIEW: I tested it and Qwen 3.8 benchmarks, published for a model with no public benchmarks and no general API. Hands-on reviews of software nobody outside a preview can access are a useful reminder to check whether a page's numbers have a source you can click.

Test claims against your own tasks, not against announcements

DataLLM Lab gives you one OpenAI-compatible key for 300+ models at https://www.datallmlab.com/v1. Point your existing eval harness at several models, swap only the model string, and rank them on your workload — so that the day Qwen3.8 opens up, you already know what score it has to beat.

What a well-evidenced launch looks like

Three days before WAIC, on 16 July 2026, Moonshot AI released Kimi K3 — a 2.8-trillion-parameter MoE with a 1M-token context, priced at $3.00 per million input and $15.00 per million output. Within days it was independently measured: 57 on the Artificial Analysis Intelligence Index, rank #4 of 186, with Artificial Analysis also noting the trade-offs (slow at 36.9 tokens/sec, and very verbose). It was also reported as #1 on the blind Frontend Code Arena, ahead of Claude Fable 5 — we could not verify that leaderboard directly, so treat it as reported by Tom's Hardware and other outlets rather than as first-hand.

One correction to how K3 is often described: it was not an open-weight launch. It shipped API-only on 16 July; weights are dated 27 July 2026 under an expected Modified MIT license, and as of 21 July the newest model on Moonshot's Hugging Face organisation is still K2.7 Code. Open-weights-dated, not open-weight. We cover it in full in our Kimi K3 review.

This is not a story about which country's labs overclaim. K3 is from a Chinese lab too, launched three days earlier, and it is the example of doing this right. The variable is evidence standards, not origin.

Evidence you can act on Qwen3.8-Max-Preview Kimi K3 Qwen3.7-Max
Announced 19 Jul 2026 (preview) 16 Jul 2026 Shipped, GA
Independent index score None — not on the board 57 (#4 of 186) 46 (#27 of 186)
Published per-token price None from Alibaba — Credits only $3.00 / $15.00 per M $2.50 / $7.50 per M
Context window ~1M, unconfirmed on vendor docs 1M 1M
Blind human evaluation None Reported #1, Frontend Code Arena Not applicable here
Weights Soon, no date, no license Dated 27 Jul 2026, not yet on HF API-only, proprietary
Can you budget with it today? No Yes Yes

An open-weight Max release would genuinely be a departure — Alibaba has historically open-sourced the smaller and mid-tier Qwen models while keeping the flagship Max line proprietary and API-only, which is exactly the state Qwen3.7-Max is in today. If it happens, it changes the open-weight landscape materially; see our running best open-source LLM in 2026 comparison and the Qwen vs DeepSeek head-to-head for where the open tiers currently sit.

What would settle it

The nice thing about a numeric claim is that it is falsifiable. Two events would resolve this within days of general access:

  1. An Artificial Analysis Intelligence Index entry above 57. That clears Kimi K3. Above 59 clears GPT-5.6 Sol and makes the literal claim of second place true. Below 57 does not.
  2. A blind arena placement above Claude Fable 5. Human preference under blind conditions is the check that vendor-selected benchmarks cannot game.

Two dates to watch: 27 July 2026, the date on Kimi K3's promised weights, and the Qwen3.8 open-weight release, which currently has no date at all. A third signal, easy to miss: the day qwen3.8-max-preview appears on the international Token Plan model table with a per-token rate, it becomes a product you can plan around rather than a preview you can only sample.

A decision rule for teams

If you are deciding what to do this week, the rule is short:

None of this is a prediction that Qwen3.8 will disappoint. Qwen has shipped genuinely strong models, and the coding and agentic direction Alibaba describes is plausible. The point is narrower and, we think, more useful: as of 21 July 2026, the correct amount of confidence to have in the second-only-to-Fable-5 claim is zero — not because it is implausible, but because nothing has been published that could make it more than zero. Come back when there is a number.

FAQ

Is Qwen 3.8 Max actually better than Kimi K3?

There is no evidence either way. As of 21 July 2026 Qwen3.8-Max does not appear anywhere on the Artificial Analysis Intelligence Index leaderboard, which lists 186 models — Qwen entries stop at Qwen3.7 Max. Kimi K3 does appear, scoring 57 and ranking #4. Alibaba's ranking claim is vendor-reported and currently untestable by anyone outside the preview.

How many parameters does Qwen 3.8 Max have?

Alibaba states 2.4 trillion total parameters. That figure is vendor-reported and unverified — there is no model card to check it against. More important for cost and latency estimates: the activated-parameter count per token has not been disclosed at all, and the sparse Mixture-of-Experts architecture is inferred from previous Qwen Max models rather than confirmed in the announcement.

What does Qwen 3.8 Max Preview cost per million tokens?

Alibaba has not published a per-million-token rate for it. Preview access runs through the Token Plan subscription on Model Studio and Qwen Cloud, plus the Qoder and QoderWork platforms, billed in Credits: Lite 60 CNY/month (limited-time 39), Standard 180 (139), Pro 600 (499), with a promotion consuming Credits at as low as 10% of a standard rate that has never been published for this model. Third-party gateways do list per-token rates — NanoGPT at $1.50 / $5.00 per million, AIHubMix at $0.17 / $0.51 — but they disagree by roughly 9x and neither is Alibaba's own rate card.

Is Qwen 3.8 open weight?

Not yet. The Qwen team said it is going open-weight soon, with no date, no license and no Hugging Face repository as of 21 July 2026. It would be a departure from Alibaba's pattern of open-sourcing smaller and mid-tier Qwen models while keeping the flagship Max line proprietary and API-only, as Qwen3.7-Max still is.

What is the Qwen 3.8 Max context window?

Roughly 1 million tokens, directionally. A precise figure of 983,616 tokens with 131,072 max output circulates from integration metadata, but it does not appear on any Alibaba Cloud or Qwen Cloud documentation page we could fetch — both omit context specifications entirely. Treat the exact number as unconfirmed.

Should I plan a migration to Qwen 3.8 Max now?

No. There is no first-party unit price, so you cannot build a defensible cost per task; there is no general availability outside the preview surfaces; and there is no independent quality measurement to justify a switch. Keep your evaluation harness ready, watch for an Artificial Analysis entry or a blind arena placement, and run your own tasks the day general access opens.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.