Ox Alpha Was GLM-5.3-Flash — and Our Measurement Pointed the Wrong Way
On 2026-08-26 Z.ai confirmed that Ox Alpha was GLM-5.3-Flash. Six days earlier we published a measurement-based argument that the GLM guess was the one our data contradicted most directly. That conclusion was wrong. The interesting part is why, because it is not a mistake anyone can avoid by measuring harder. We ran the model on our nine executed Python tasks under both names. Under Ox Alpha it scored 9 of 9 and reported 0 reasoning tokens across all nine. Under GLM-5.3-Flash it scored 9 of 9 and reported 1,212, on 3.1x the output tokens. Same model, same tasks, same harness. The deployment configures the behaviour, so behaviour cannot identify the model — which is precisely the caveat we flagged as the biggest hole in our own argument, and then did not weight heavily enough.
Correction, October 3, 2026: single-model benchmark latency is a mean, not a median. Labels were corrected; measured values and the original run dates are unchanged. Measurement details.
This page used to argue that Ox Alpha was not a GLM model. It was. Rather than quietly swap the conclusion, here is the whole thing: what we measured, why it pointed the wrong way, and the follow-up measurement that explains it. The correction is more useful than the original claim was.
The reveal
- 2026-08-20 — a listing called
stealth/ox-alphaappeared on OpenRouter: free, no owner, no model card, a 1M-token context window. - 2026-08-22 — we ran it on our benchmark and published our analysis.
- 2026-08-26 — Bloomberg reported Z.ai's confirmation that Ox Alpha was a new model in the GLM series. It is GLM-5.3-Flash. By that point it was reported as the most-used model on OpenRouter, having processed roughly 23 trillion tokens in six days.
- 2026-08-31 —
stealth/ox-alphais gone.z-ai/glm-5.3-flashis live, so we re-ran the identical suite.
The same model, measured twice
Both runs: same nine tasks, same prompts, temperature 0, max_tokens 4000, one scored attempt each.
as stealth/ox-alpha | as z-ai/glm-5.3-flash | Change | |
|---|---|---|---|
| Measured | 2026-08-22 | 2026-08-31 | — |
| Score | 9/9 | 9/9 | none |
| Reasoning tokens | 0 | 1,212 | 0 to 1,212 |
| Output tokens | 3,833 | 11,884 | 3.1x |
| Input tokens | 1,330 | 655 | −51% |
| Mean latency | 17.6s | 24.9s | +41% |
| Price | free (promotional) | $0.075 / $0.25 | — |
Nothing about the weights changed between those two rows. What changed is how the endpoint was configured to run them.
What we got wrong, precisely
Two errors, and they are different sizes.
The small one: we tested the wrong family member. The community guess was “GLM-5.3 Flash”. We benchmarked z-ai/glm-5.3 — a different model, 744B total and 40B active against Flash's 320B and 18B. We did note in the limits section that Flash was a separate variant we had not tested. But we then let “GLM-5.3 does not match” stand as a rebuttal of a guess about Flash. Testing a sibling and reporting it against the named model is a real error, and the limits note did not license the headline.
The large one: we treated a configurable behaviour as an intrinsic property. Our whole argument rested on Ox Alpha reporting zero reasoning tokens while every GLM generation reports hundreds. We even probed the raw usage payload to confirm the field was populated rather than missing — and it was, at zero. That measurement was correct. The inference from it was not, because a provider can serve the same weights with thinking on or off, and the stealth endpoint had it off.
What the method got right
Not everything failed, and it matters which parts held.
- The measurements were accurate. Ox Alpha really did report 0 reasoning tokens on all nine tasks. The probe that confirmed the field was populated rather than absent was the right check to run, and it returned the right answer.
- We refused to name a lab. The original article said plainly: “It does not identify the model. Ruling out a match on one behavioural axis is not the same as naming the lab.” We did not claim to have solved it, and that restraint is the only reason this correction is small rather than embarrassing.
- We wrote down the exact hole that sank us. Verbatim from the original: “A stealth endpoint could be deliberately set to route reasoning into visible content, which would make a heavy-reasoning model look like a zero-reasoning one. We cannot exclude that, and it is the single biggest hole in this argument.” That was right. The lesson is not that we failed to think of it — it is that we listed it as a caveat and still let the headline lean the other way.
A caveat you publish and then argue past is not a caveat. That is the part we are changing.
Why fingerprinting a stealth model cannot work
The generalisable point is worth more than the specific answer.
Everything you can observe from outside an endpoint — reasoning-token counts, latency, verbosity, refusal style, formatting habits — is downstream of serving configuration, not just weights. Thinking budgets, system prompts, sampling parameters and reasoning visibility are all set by whoever runs the endpoint. A lab shipping a model anonymously has every reason to run it in an unusual configuration, deliberately or otherwise.
So a behavioural mismatch cannot rule a candidate out, and a behavioural match cannot rule one in. In hindsight there was a signal we noticed and underweighted: Ox Alpha's prose style. Asked to work through 17 × 23, it split the problem as 17 × 20 plus 17 × 3 under a bolded heading. Asked the same question today, GLM-5.3-Flash produces the same decomposition under the same formatting. Style survived the configuration change; token accounting did not. We weighted the hard-looking number over the soft-looking one, and the soft one was the load-bearing signal.
This is the second time our own harness has taught us that an unusual number needs a mechanism before it needs a conclusion. The first was a content filter that made a strong model look like it failed five tasks.
What GLM-5.3-Flash actually is
Third-party specifications, read 2026-08-31:
- 320B total parameters, 18B active per token, and the first natively multimodal model in the GLM-5 series.
- A hybrid sparse and linear attention architecture, with a 1,310,720-token context window.
- Weights on Hugging Face under the MIT licence.
- Launch promotional pricing of $0.075 in and $0.25 out per million tokens, doubling to $0.15 and $0.50 after 2026-09-09.
Our measured figure of $0.34 per 1,000 tasks uses the promotional price captured 2026-08-31 and will roughly double when that promotion ends — a good illustration of why every price needs a date attached. Even at the post-promotional rate it is inexpensive: at $0.34 it would rank fifth cheapest among the 45 models in our set that scored 9 out of 9, behind DeepSeek V3.2 at $0.08, Qwen3 Coder Next and DeepSeek Chat at $0.10, and DeepSeek V4-Flash at $0.13.
For the larger sibling we originally tested, our GLM-5.3 review has the numbers — and it needs its own correction, since it repeated this page's conclusion. It now points here.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Both runs used the identical prompts. Cost is derived from measured token counts at the list price on the measurement date, not a billing statement. Runs go through OpenRouter. Full method on the methodology page. Our benchmark data records stealth/ox-alpha with a revealedAs field pointing at z-ai/glm-5.3-flash, so the two rows cannot drift apart again.
FAQ
What was Ox Alpha? GLM-5.3-Flash, from Z.ai. Confirmed 2026-08-26, six days after the stealth listing appeared.
Did you correctly identify it? No. We argued from reasoning-token behaviour that the GLM guess was the least likely, and the GLM guess was right.
Why did it report zero reasoning tokens under the stealth name? The stealth endpoint served the model with thinking suppressed. Under its own name, on the identical tasks, it reported 1,212 reasoning tokens and emitted 3.1x the output.
Was the benchmark itself wrong? No. Both runs are accurate; they measure two different endpoint configurations of one model. The error was the inference, not the instrument.
Can you identify a stealth model from its behaviour? No, and this page is the counter-example. Observable behaviour reflects serving configuration as much as weights.
Is GLM-5.3-Flash worth using? On our tasks, 9 out of 9 at $0.34 per 1,000 at promotional pricing, which would place it fifth cheapest among our 9/9 models. The promotion ends 2026-09-09.
DataLLM Lab