Benchmarks

Ox Alpha Was GLM-5.3-Flash — and Our Measurement Pointed the Wrong Way

On 2026-08-26 Z.ai confirmed that Ox Alpha was GLM-5.3-Flash. Six days earlier we published a measurement-based argument that the GLM guess was the one our data contradicted most directly. That conclusion was wrong. The interesting part is why, because it is not a mistake anyone can avoid by measuring harder. We ran the model on our nine executed Python tasks under both names. Under Ox Alpha it scored 9 of 9 and reported 0 reasoning tokens across all nine. Under GLM-5.3-Flash it scored 9 of 9 and reported 1,212, on 3.1x the output tokens. Same model, same tasks, same harness. The deployment configures the behaviour, so behaviour cannot identify the model — which is precisely the caveat we flagged as the biggest hole in our own argument, and then did not weight heavily enough.

Correction, October 3, 2026: single-model benchmark latency is a mean, not a median. Labels were corrected; measured values and the original run dates are unchanged. Measurement details.

Chart comparing output and reasoning tokens for the same model measured under the Ox Alpha stealth name and under the GLM-5.3-Flash name

This page used to argue that Ox Alpha was not a GLM model. It was. Rather than quietly swap the conclusion, here is the whole thing: what we measured, why it pointed the wrong way, and the follow-up measurement that explains it. The correction is more useful than the original claim was.

The reveal

The same model, measured twice

Both runs: same nine tasks, same prompts, temperature 0, max_tokens 4000, one scored attempt each.

as stealth/ox-alphaas z-ai/glm-5.3-flashChange
Measured2026-08-222026-08-31—
Score9/99/9none
Reasoning tokens01,2120 to 1,212
Output tokens3,83311,8843.1x
Input tokens1,330655−51%
Mean latency17.6s24.9s+41%
Pricefree (promotional)$0.075 / $0.25—
One model, two names, two completely different token profilesIdentical nine tasks, identical prompts, temperature 0. Both runs scored 9/9.as stealth/ox-alpha2026-08-223,833 output · 0 reasoningas z-ai/glm-5.3-flash2026-08-311,21211,884Blue = reasoning tokens. Grey = the rest of the output. One scale: 0.0454 px per token.The stealth endpoint was serving the model with its thinking suppressed. Nothing about the weights changed between these two rows.
The behaviour we fingerprinted was a property of the endpoint, not of the model.

Nothing about the weights changed between those two rows. What changed is how the endpoint was configured to run them.

What we got wrong, precisely

Two errors, and they are different sizes.

The small one: we tested the wrong family member. The community guess was “GLM-5.3 Flash”. We benchmarked z-ai/glm-5.3 — a different model, 744B total and 40B active against Flash's 320B and 18B. We did note in the limits section that Flash was a separate variant we had not tested. But we then let “GLM-5.3 does not match” stand as a rebuttal of a guess about Flash. Testing a sibling and reporting it against the named model is a real error, and the limits note did not license the headline.

The large one: we treated a configurable behaviour as an intrinsic property. Our whole argument rested on Ox Alpha reporting zero reasoning tokens while every GLM generation reports hundreds. We even probed the raw usage payload to confirm the field was populated rather than missing — and it was, at zero. That measurement was correct. The inference from it was not, because a provider can serve the same weights with thinking on or off, and the stealth endpoint had it off.

Not everything failed, and it matters which parts held.

A caveat you publish and then argue past is not a caveat. That is the part we are changing.

Why fingerprinting a stealth model cannot work

The generalisable point is worth more than the specific answer.

Everything you can observe from outside an endpoint — reasoning-token counts, latency, verbosity, refusal style, formatting habits — is downstream of serving configuration, not just weights. Thinking budgets, system prompts, sampling parameters and reasoning visibility are all set by whoever runs the endpoint. A lab shipping a model anonymously has every reason to run it in an unusual configuration, deliberately or otherwise.

So a behavioural mismatch cannot rule a candidate out, and a behavioural match cannot rule one in. In hindsight there was a signal we noticed and underweighted: Ox Alpha's prose style. Asked to work through 17 × 23, it split the problem as 17 × 20 plus 17 × 3 under a bolded heading. Asked the same question today, GLM-5.3-Flash produces the same decomposition under the same formatting. Style survived the configuration change; token accounting did not. We weighted the hard-looking number over the soft-looking one, and the soft one was the load-bearing signal.

This is the second time our own harness has taught us that an unusual number needs a mechanism before it needs a conclusion. The first was a content filter that made a strong model look like it failed five tasks.

What GLM-5.3-Flash actually is

Third-party specifications, read 2026-08-31:

Our measured figure of $0.34 per 1,000 tasks uses the promotional price captured 2026-08-31 and will roughly double when that promotion ends — a good illustration of why every price needs a date attached. Even at the post-promotional rate it is inexpensive: at $0.34 it would rank fifth cheapest among the 45 models in our set that scored 9 out of 9, behind DeepSeek V3.2 at $0.08, Qwen3 Coder Next and DeepSeek Chat at $0.10, and DeepSeek V4-Flash at $0.13.

For the larger sibling we originally tested, our GLM-5.3 review has the numbers — and it needs its own correction, since it repeated this page's conclusion. It now points here.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Both runs used the identical prompts. Cost is derived from measured token counts at the list price on the measurement date, not a billing statement. Runs go through OpenRouter. Full method on the methodology page. Our benchmark data records stealth/ox-alpha with a revealedAs field pointing at z-ai/glm-5.3-flash, so the two rows cannot drift apart again.

FAQ

What was Ox Alpha? GLM-5.3-Flash, from Z.ai. Confirmed 2026-08-26, six days after the stealth listing appeared.

Did you correctly identify it? No. We argued from reasoning-token behaviour that the GLM guess was the least likely, and the GLM guess was right.

Why did it report zero reasoning tokens under the stealth name? The stealth endpoint served the model with thinking suppressed. Under its own name, on the identical tasks, it reported 1,212 reasoning tokens and emitted 3.1x the output.

Was the benchmark itself wrong? No. Both runs are accurate; they measure two different endpoint configurations of one model. The error was the inference, not the instrument.

Can you identify a stealth model from its behaviour? No, and this page is the counter-example. Observable behaviour reflects serving configuration as much as weights.

Is GLM-5.3-Flash worth using? On our tasks, 9 out of 9 at $0.34 per 1,000 at promotional pricing, which would place it fifth cheapest among our 9/9 models. The promotion ends 2026-09-09.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.