Open Models

Tencent Hy4 Preview: A 0.07-Point Quality Claim and a 3x Cost Spread

Tencent open-sourced Hy4 Preview on 2026-08-28 under Apache 2.0: 770B total parameters, 49B active per token, a 1M-token context window. Its headline evidence is an internal blind evaluation — 163 experts, 203 engineering tasks — scoring Hy4 at 2.99 out of 4 against GLM-5.3 at 2.92 and Kimi K3 at 2.94. Those are margins of 0.07 and 0.05 points on a four-point scale. We have measured all three models on our own executed benchmark, and the interesting comparison is not the quality gap, which is small and self-reported. It is the cost gap, which is large and checkable: Hy4 Preview $2.08, GLM-5.3 $3.35, Kimi K3 $6.33 per 1,000 tasks. One caveat up front, and it is not a small one: our Hy4 run did not complete.

Correction, October 3, 2026: single-model benchmark latency is a mean, not a median. Labels were corrected; measured values and the original run dates are unchanged. Measurement details.

Chart comparing measured cost per 1,000 tasks for Hy4 Preview against GLM-5.3 and Kimi K3

A vendor blind evaluation that names its competitors is more useful than most benchmark tables, because it tells you exactly who the vendor thinks it is beating. Tencent named GLM-5.3 and Kimi K3. We have run both.

What Tencent shipped

Third-party, from Tencent and the release coverage, read 2026-08-31:

Our run, including the task it failed

We ran Hy4 Preview on our nine executed Python tasks. It solved eight and then failed the ninth at the API layer — parse_csv_line returned an empty body with finish_reason=length after retries, meaning it spent the entire 4,000-token allowance on reasoning without emitting an answer.

TaskResultLatencyOutput tokensReasoning tokens
two_sumpass10.7s397340
valid_parenthesespass6.2s259179
merge_intervalspass10.4s452355
roman_to_intpass19.8s1,050930
lcs_lenpass6.9s329235
flattenpass21.9s1,2131,130
top_k_wordspass10.4s521438
token_bucketpass42.3s2,2172,101
parse_csv_lineno answer———
Total8 of 8 scored16.1s mean—714 per call

We have marked Hy4 Preview excluded in our data rather than giving it a score. Printing “8/8” in a column where every other model shows “9/9” implies a comparison that does not exist, because the denominators differ. The $2.08 figure below is measured on the eight tasks it did complete and is therefore an underestimate of a full run.

That failure mode is not unique to Hy4 — we have now seen it on several models, and it is what Step 3.5 Flash does on almost half our suite. It is worth knowing about before you set a token cap in production.

The quality claim next to the cost

A 0.07-point quality spread against a 3x cost spreadLeft: Tencent’s own blind evaluation, out of 4. Right: our measured cost per 1,000 tasks.TENCENT’S BLIND EVAL (of 4)Hy4 Preview2.99Kimi K32.94GLM-5.32.92OUR MEASURED COST ($ / 1k tasks)Hy4 Preview$2.08GLM-5.3$3.35Kimi K3$6.33Left bars start at 2.80 so the 0.07-point spread is visible at all; right bars start at zero. The scales are not comparable —that is the point. Hy4’s $2.08 is measured on 8 of 9 tasks, so it understates a complete run.Quality figures are Tencent’s, self-reported. Cost figures are ours, from executed code, priced on each model’s measurement date.
The models Tencent chose to compare against cost 1.6x and 3x more on our tasks.

Read charitably, this is a good result for Tencent: a self-reported dead heat on quality against models that cost substantially more to run. Read sceptically, a 0.07-point margin on a vendor's own blind evaluation is inside the noise of any evaluation we have ever seen published, and the honest reading of 2.99 against 2.94 is these three models are the same.

Either reading points the same way. If the quality is a tie and the price is not, price decides. And on price, none of these three is the answer — GLM-5.3-Flash scored 9 of 9 at $0.34 and DeepSeek V3.2 at $0.08, both far below the cheapest model in Tencent's comparison.

The architecture, briefly

Two details worth noting, because they are unusual rather than because we tested them. The attention module uses Gated DeepSeek Sparse Attention with an IndexCache that reuses sparse indices across layers — an inference-cost optimisation, not a quality one. And the expert layout is asymmetric: layer 1 is a dense feed-forward network while layers 2 through 78 route to 8 of 256 experts plus a shared expert.

At 770B total parameters, self-hosting is a datacentre decision even with the FP8 checkpoint. The Apache 2.0 licence is the genuinely permissive part — see open weights versus open source for what that does and does not grant.

Where it fits

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived from measured token counts at the list price captured 2026-08-31, not a billing statement, and prices move. Runs go through OpenRouter. Full method on the methodology page.

What we did not measure

FAQ

What is Tencent Hy4 Preview? An Apache 2.0 open-weight mixture-of-experts model released 2026-08-28, with 770B total parameters, 49B active per token and a 1M-token context window.

Is Hy4 Preview better than GLM-5.3? Tencent's own blind evaluation says 2.99 against 2.92 out of 4. That is a 0.07-point margin, self-reported, and we would not call it a difference.

How much does Hy4 Preview cost? $0.83 per million input tokens and $2.50 output as of 2026-08-31. On the eight tasks it completed, that came to $2.08 per 1,000 tasks.

Did Hy4 Preview pass your benchmark? Not completely. It answered eight of nine tasks; the ninth exhausted its token budget on reasoning and returned nothing, so we excluded it rather than publish a partial score.

What licence is Hy4 Preview under? Apache 2.0.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.