GLM-5.3 Review: The Version Number Does Not Predict the Bill
GLM-5.3 scored 9 out of 9 on our executed Python benchmark at $3.35 per 1,000 tasks and a 14.2-second mean. GLM-5.2 scored the same 9 out of 9 at $1.99 — so the newer release costs 68% more for an identical result, and it is slower. That is not a one-off. We have now run four GLM generations on the same nine tasks, and all four scored 9 out of 9 while their measured cost bounced between $1.99 and $4.35 with no relation to the version number. GLM-5.3 is also the model most people guessed the free stealth endpoint Ox Alpha was. We argued from measurement that it was not — and we were wrong. Ox Alpha turned out to be GLM-5.3-Flash, a smaller sibling of the model reviewed here. The correction, and why our reasoning-token evidence pointed the wrong way, is here.
Correction, October 3, 2026: single-model benchmark latency is a mean, not a median. Labels were corrected; measured values and the original run dates are unchanged. Measurement details.
Z.ai ships GLM point releases fast enough that most coverage treats the newest as the best by default. Four generations through the same harness says otherwise.
The result
| Metric | GLM-5.3 |
|---|---|
| Score | 9/9 |
| Measured cost / 1,000 tasks | $3.35 |
| Mean latency | 14.2s |
| Reasoning tokens per call | 631 |
| Output tokens on the suite | 6,640 |
| List price in / out | $1.40 / $4.40 per 1M |
| Context window | 1,048,576 tokens |
| Measured | 2026-08-22 |
5.3 against 5.2
| GLM-5.2 | GLM-5.3 | Change | |
|---|---|---|---|
| List price in / out | $0.97 / $3.04 | $1.40 / $4.40 | +44% |
| Context window | 1,048,576 | 1,048,576 | none |
| Score | 9/9 | 9/9 | none |
| Measured cost / 1k tasks | $1.99 | $3.35 | +68% |
| Mean latency | 12.3s | 14.2s | +15% |
| Reasoning tokens | 559 | 631 | +13% |
The list price went up 44% and the bill went up 68%, because the model also got slightly more talkative. Same window, same score. On this suite there is nothing 5.3 does that 5.2 does not.
Four generations, one score, four prices
GLM-5.1 was the expensive outlier at $4.35 and 23.6 seconds on 1,327 reasoning tokens. GLM-5.2 cut that to $1.99 and 12.3 seconds. GLM-5.3 gives about half of that saving back. Every one of the four solved all nine tasks, so on this workload the version number tells you nothing about quality and quite a lot about price — in an unhelpful direction.
We keep measuring the same thing across vendors. Grok 4.6 costs 70% more than Grok 4.5 at an identical list price. GPT-5.1-Codex-Max lists below GPT-5.2-Codex and costs 3.09x more. Qwen3.8-Max raised its price 36% and cut the bill anyway. Price per token is half the equation; token volume is the half nobody publishes.
The Ox Alpha question, and our mistake
When the free stealth endpoint stealth/ox-alpha appeared, the most repeated guess was GLM-5.3 Flash. We ran this model — z-ai/glm-5.3, not Flash — on the same suite the same day, found 631 reasoning tokens against Ox Alpha's zero, and treated that as evidence against the GLM guess.
Z.ai confirmed on 2026-08-26 that Ox Alpha was GLM-5.3-Flash. The guess was right and our rebuttal was wrong, for two reasons: Flash is a different model from the one on this page (320B total and 18B active, against 744B and 40B), and reasoning-token behaviour is set by the endpoint rather than the weights. Re-running the model under its own name on the identical tasks produced 1,212 reasoning tokens where the stealth endpoint produced 0. The full correction is here.
Which GLM to run
- GLM-5.2 for coding of this shape: cheapest of the four, same 9/9, same million-token window as 5.3, and faster.
- GLM-5.3 if you have measured a workload where it beats 5.2. We could not find one on nine self-contained Python functions, and it is 68% more expensive.
- GLM-5 or 5.1 only if you are already pinned to them. Both carry a 204,800-token window, a fifth of what 5.2 and 5.3 offer.
Against the wider open field GLM is mid-priced: DeepSeek V3.2 scored 9/9 at $0.08 and Qwen3 Coder Next at $0.10. GLM-5.3 ranks 28th cheapest among the 44 models in our set that scored 9 out of 9. If you are running GLM through a coding agent, the GLM Coding Plan page covers the subscription route and using Claude Code with GLM covers the wiring.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived — measured token counts multiplied by the list price on the date shown, not a billing statement. The four generations were measured on different dates at their then-current prices, which is why every figure needs a date. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- The million-token context. Our prompts are short; nothing here tests long-context recall.
- Self-hosting. We called the hosted endpoint. GLM ships open weights, and the economics of running them yourself are a different calculation entirely — see the open-source guide.
- Anything but Python, and only nine self-contained functions.
- Repeat runs. One scored attempt per task. Single-run figures, not averages.
- GLM-5.3 Flash, a separate variant, is not what we tested. We tested
z-ai/glm-5.3.
FAQ
Is GLM-5.3 better than GLM-5.2? Not on our nine tasks. Both scored 9 out of 9; 5.3 cost 68% more and was 15% slower.
How much does GLM-5.3 cost? It lists at $1.40 per million input tokens and $4.40 output as of 2026-08-22. On our suite that came to $3.35 per 1,000 tasks.
Which GLM is cheapest per finished task? Of the four we have measured, GLM-5.2 at $1.99.
Was Ox Alpha GLM-5.3? It was GLM-5.3-Flash, a smaller sibling of the model on this page, confirmed by Z.ai on 2026-08-26. We originally argued against the GLM guess and were wrong.
What is GLM-5.3's context window? 1,048,576 tokens, the same as GLM-5.2 and five times GLM-5 and 5.1.
DataLLM Lab