GLM-5.3-Flash Review: A Tenth the Cost of GLM-5.3, and Slower Than All of Them
GLM-5.3-Flash scored 9 out of 9 on our executed Python benchmark at $0.34 per 1,000 tasks. The full-size GLM-5.3 scored the same at $3.35 — so Flash costs a tenth as much for an identical result. It gets there the hard way: at 24.9 seconds it is the slowest of the five GLM generations we have measured, and it emitted 11,884 output tokens against GLM-5.3's 6,640. Nearly twice the tokens, a tenth of the bill, because the list price is $0.075 and $0.25 against $1.40 and $4.40. This is the case where price per token beats volume outright — the opposite of what we have been finding all month. You may also know this model by another name: it shipped anonymously as Ox Alpha first.
Correction, October 3, 2026: single-model benchmark latency is a mean, not a median. Labels were corrected; measured values and the original run dates are unchanged. Measurement details.
Z.ai shipped this model twice: once anonymously in August as a free stealth endpoint, and once under its own name at a promotional price. We measured it both times. This page is about the named version, which is the one you can plan around.
The result
| Metric | GLM-5.3-Flash |
|---|---|
| Score | 9/9 |
| Measured cost / 1,000 tasks | $0.34 |
| Mean latency | 24.9s |
| Reasoning tokens per call | 1,212 |
| Output tokens on the suite | 11,884 |
| List price in / out | $0.075 / $0.25 per 1M (promotional) |
| Context window | 1,310,720 tokens |
| Measured | 2026-08-31 |
Per task, it ranged from 388 output tokens on flatten to 2,492 on lcs_len, and from 7.9 seconds to 43.3. It is not a consistent model; it is a cheap one that finishes.
Against the rest of the GLM family
Five generations, identical nine tasks, all scoring 9 out of 9:
| Model | List price in / out | Measured cost / 1k | Latency | Reasoning tokens | Context |
|---|---|---|---|---|---|
| GLM-5.3-Flash | $0.075 / $0.25 | $0.34 | 24.9s | 1,212 | 1,310,720 |
| GLM-5.2 | $1.19 / $3.74 | $1.99 | 12.3s | 559 | 1,048,576 |
| GLM-5 | $0.60 / $1.92 | $2.30 | 16.5s | 719 | 204,800 |
| GLM-5.3 | $1.40 / $4.40 | $3.35 | 14.2s | 631 | 1,310,720 |
| GLM-5.1 | $0.97 / $3.04 | $4.35 | 23.6s | 1,327 | 204,800 |
Cheapest and slowest at the same time
The usual trade is that a cheap model is fast because it thinks less. Flash inverts it. It is the slowest of the five at 24.9 seconds and emits more reasoning than every generation except GLM-5.1. It is cheap anyway, because $0.25 per million output tokens is a seventeenth of GLM-5.3's $4.40.
That is worth pausing on, because it runs against everything else we measured this month. GPT-5.1-Codex-Max lists below its sibling and costs 3.09x more, because volume swamped price. Grok 4.6 matches Grok 4.5's sticker exactly and costs 70% more, same reason. Here the arrow points the other way: Flash emits 1.8x GLM-5.3's output tokens and still bills a tenth, because the price gap is far larger than the volume gap.
Neither factor wins by default. That is the whole argument for measuring cost per finished task instead of reading a price page or a token count in isolation.
The price is promotional until 2026-09-09
The $0.075 and $0.25 we measured against is launch pricing. Z.ai has published that it rises to $0.15 and $0.50 on 2026-09-09, doubling both sides. Our $0.34 per 1,000 tasks will roughly double with it, to somewhere near $0.68 on the same token counts.
Even then it stays the cheapest GLM by a wide margin. But every figure on this page carries a capture date of 2026-08-31 for exactly this reason, and if you are reading this after 2026-09-09 the price you see will not be the price we measured. That is the general case, not a special one — see why every LLM price needs a date attached.
What GLM-5.3-Flash is
Third-party specifications, read 2026-08-31:
- 320B total parameters, 18B active per token.
- The first natively multimodal model in the GLM-5 series, with a hybrid sparse and linear attention architecture.
- A 1,310,720-token context window, the same as full-size GLM-5.3 and larger than GLM-5.2's.
- Weights published to Hugging Face under the MIT licence.
- It shipped first as an unnamed free endpoint,
stealth/ox-alpha, on 2026-08-20, and Z.ai confirmed the identity on 2026-08-26. We measured it under both names and got very different token profiles — worth reading before you trust any behavioural claim about this model.
Who should run it
- Batch and background work: yes. 24.9 seconds does not matter in a queue, and a tenth of the price does.
- Interactive tools: probably not. GLM-5.2 is twice as fast for six times the money, and GPT-5.4 mini scored the same 9/9 in 2.3 seconds at $0.53.
- Long context: it has the largest window in the family at 1,310,720 tokens, tied with GLM-5.3 and at a tenth of the cost.
- The absolute cheapest: still not this. DeepSeek V3.2 scored 9/9 at $0.08 in 7.1 seconds. GLM-5.3-Flash ranks fifth cheapest among the 45 models in our set that scored 9 out of 9.
Our open-source guide covers the wider field, and the GLM Coding Plan page covers the subscription route if you are running GLM through an agent.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived from measured token counts at the list price captured 2026-08-31, not a billing statement. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- The 1.3M context window. Our prompts are short; nothing here tests long-context recall.
- Multimodal input, which is the headline architectural claim. Text only.
- Self-hosting. Weights are MIT, and running them yourself is a different calculation.
- Post-promotional pricing. We measured at launch pricing. The 2026-09-09 figures are Z.ai's published plan, not something we have billed against.
- Repeat runs. One scored attempt per task. Single-run figures, not averages.
FAQ
How much does GLM-5.3-Flash cost? $0.075 per million input tokens and $0.25 output at launch pricing, rising to $0.15 and $0.50 on 2026-09-09. On our nine tasks that worked out to $0.34 per 1,000 tasks.
Is GLM-5.3-Flash as good as GLM-5.3? On our tasks, yes — both scored 9 out of 9. Flash cost a tenth as much and took 75% longer.
How fast is GLM-5.3-Flash? 24.9 seconds mean, the slowest of the five GLM models we have measured.
Was GLM-5.3-Flash the Ox Alpha stealth model? Yes, confirmed by Z.ai on 2026-08-26.
What licence is it under? MIT, with weights on Hugging Face.
DataLLM Lab