Tencent Hy4 Preview: A 0.07-Point Quality Claim and a 3x Cost Spread
Tencent open-sourced Hy4 Preview on 2026-08-28 under Apache 2.0: 770B total parameters, 49B active per token, a 1M-token context window. Its headline evidence is an internal blind evaluation — 163 experts, 203 engineering tasks — scoring Hy4 at 2.99 out of 4 against GLM-5.3 at 2.92 and Kimi K3 at 2.94. Those are margins of 0.07 and 0.05 points on a four-point scale. We have measured all three models on our own executed benchmark, and the interesting comparison is not the quality gap, which is small and self-reported. It is the cost gap, which is large and checkable: Hy4 Preview $2.08, GLM-5.3 $3.35, Kimi K3 $6.33 per 1,000 tasks. One caveat up front, and it is not a small one: our Hy4 run did not complete.
Correction, October 3, 2026: single-model benchmark latency is a mean, not a median. Labels were corrected; measured values and the original run dates are unchanged. Measurement details.
A vendor blind evaluation that names its competitors is more useful than most benchmark tables, because it tells you exactly who the vendor thinks it is beating. Tencent named GLM-5.3 and Kimi K3. We have run both.
What Tencent shipped
Third-party, from Tencent and the release coverage, read 2026-08-31:
- Released and open-sourced 2026-08-28, under Apache 2.0, repository
Tencent-Hunyuan/Hy4-preview. - 770B total parameters, 49B activated per token, across 78 layers. The first layer uses a dense feed-forward network; the other 77 use 256 routed experts with top-8 routing plus one shared expert.
- A 1,000,000-token context window, with an FP8 quantised variant shipped alongside.
- Available through WorkBuddy, CodeBuddy, Yuanbao and ima, with API access via Tencent Cloud TokenHub and OpenRouter.
- The evaluation: an internal blind test with 163 experts across 203 engineering tasks. Hy4 Preview 2.99 out of 4; GLM-5.3 2.92; Kimi K3 2.94.
Our run, including the task it failed
We ran Hy4 Preview on our nine executed Python tasks. It solved eight and then failed the ninth at the API layer — parse_csv_line returned an empty body with finish_reason=length after retries, meaning it spent the entire 4,000-token allowance on reasoning without emitting an answer.
| Task | Result | Latency | Output tokens | Reasoning tokens |
|---|---|---|---|---|
| two_sum | pass | 10.7s | 397 | 340 |
| valid_parentheses | pass | 6.2s | 259 | 179 |
| merge_intervals | pass | 10.4s | 452 | 355 |
| roman_to_int | pass | 19.8s | 1,050 | 930 |
| lcs_len | pass | 6.9s | 329 | 235 |
| flatten | pass | 21.9s | 1,213 | 1,130 |
| top_k_words | pass | 10.4s | 521 | 438 |
| token_bucket | pass | 42.3s | 2,217 | 2,101 |
| parse_csv_line | no answer | — | — | — |
| Total | 8 of 8 scored | 16.1s mean | — | 714 per call |
We have marked Hy4 Preview excluded in our data rather than giving it a score. Printing “8/8” in a column where every other model shows “9/9” implies a comparison that does not exist, because the denominators differ. The $2.08 figure below is measured on the eight tasks it did complete and is therefore an underestimate of a full run.
That failure mode is not unique to Hy4 — we have now seen it on several models, and it is what Step 3.5 Flash does on almost half our suite. It is worth knowing about before you set a token cap in production.
The quality claim next to the cost
Read charitably, this is a good result for Tencent: a self-reported dead heat on quality against models that cost substantially more to run. Read sceptically, a 0.07-point margin on a vendor's own blind evaluation is inside the noise of any evaluation we have ever seen published, and the honest reading of 2.99 against 2.94 is these three models are the same.
Either reading points the same way. If the quality is a tie and the price is not, price decides. And on price, none of these three is the answer — GLM-5.3-Flash scored 9 of 9 at $0.34 and DeepSeek V3.2 at $0.08, both far below the cheapest model in Tencent's comparison.
The architecture, briefly
Two details worth noting, because they are unusual rather than because we tested them. The attention module uses Gated DeepSeek Sparse Attention with an IndexCache that reuses sparse indices across layers — an inference-cost optimisation, not a quality one. And the expert layout is asymmetric: layer 1 is a dense feed-forward network while layers 2 through 78 route to 8 of 256 experts plus a shared expert.
At 770B total parameters, self-hosting is a datacentre decision even with the FP8 checkpoint. The Apache 2.0 licence is the genuinely permissive part — see open weights versus open source for what that does and does not grant.
Where it fits
- If you need Apache 2.0 weights at frontier scale, this is a serious release and the licence is the reason to care.
- If you are renting tokens, at $0.83 in and $2.50 out it is mid-priced, and our measured $2.08 per 1,000 tasks (on eight tasks) sits between GLM-5.3 and the cheap tier.
- If you are picking on cost per finished task, look lower. Our open-source guide and cheap coding roundup cover the models that clear the same suite for a fraction.
- Set your token cap deliberately. One of our nine tasks exhausted 4,000 tokens on reasoning alone.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived from measured token counts at the list price captured 2026-08-31, not a billing statement, and prices move. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- A complete run. Eight of nine tasks. The ninth never returned an answer, and our cost figure covers only the eight.
- Tencent's blind evaluation. We have not seen the 203 tasks, the rubric, or how the 163 experts were selected. Every quality figure here is Tencent's.
- The 1M context window. Our prompts are short.
- Self-hosting economics, which at 770B parameters is the main reason to care about an Apache 2.0 licence.
- Repeat runs. One scored attempt per task.
FAQ
What is Tencent Hy4 Preview? An Apache 2.0 open-weight mixture-of-experts model released 2026-08-28, with 770B total parameters, 49B active per token and a 1M-token context window.
Is Hy4 Preview better than GLM-5.3? Tencent's own blind evaluation says 2.99 against 2.92 out of 4. That is a 0.07-point margin, self-reported, and we would not call it a difference.
How much does Hy4 Preview cost? $0.83 per million input tokens and $2.50 output as of 2026-08-31. On the eight tasks it completed, that came to $2.08 per 1,000 tasks.
Did Hy4 Preview pass your benchmark? Not completely. It answered eight of nine tasks; the ninth exhausted its token budget on reasoning and returned nothing, so we excluded it rather than publish a partial score.
What licence is Hy4 Preview under? Apache 2.0.
DataLLM Lab