Model Reviews

GLM-5.3 Review: The Version Number Does Not Predict the Bill

GLM-5.3 scored 9 out of 9 on our executed Python benchmark at $3.35 per 1,000 tasks and a 14.2-second mean. GLM-5.2 scored the same 9 out of 9 at $1.99 — so the newer release costs 68% more for an identical result, and it is slower. That is not a one-off. We have now run four GLM generations on the same nine tasks, and all four scored 9 out of 9 while their measured cost bounced between $1.99 and $4.35 with no relation to the version number. GLM-5.3 is also the model most people guessed the free stealth endpoint Ox Alpha was. We argued from measurement that it was not — and we were wrong. Ox Alpha turned out to be GLM-5.3-Flash, a smaller sibling of the model reviewed here. The correction, and why our reasoning-token evidence pointed the wrong way, is here.

Correction, October 3, 2026: single-model benchmark latency is a mean, not a median. Labels were corrected; measured values and the original run dates are unchanged. Measurement details.

Chart of measured cost per 1,000 tasks across four GLM generations, all scoring 9 of 9

Z.ai ships GLM point releases fast enough that most coverage treats the newest as the best by default. Four generations through the same harness says otherwise.

The result

MetricGLM-5.3
Score9/9
Measured cost / 1,000 tasks$3.35
Mean latency14.2s
Reasoning tokens per call631
Output tokens on the suite6,640
List price in / out$1.40 / $4.40 per 1M
Context window1,048,576 tokens
Measured2026-08-22

5.3 against 5.2

GLM-5.2GLM-5.3Change
List price in / out$0.97 / $3.04$1.40 / $4.40+44%
Context window1,048,5761,048,576none
Score9/99/9none
Measured cost / 1k tasks$1.99$3.35+68%
Mean latency12.3s14.2s+15%
Reasoning tokens559631+13%

The list price went up 44% and the bill went up 68%, because the model also got slightly more talkative. Same window, same score. On this suite there is nothing 5.3 does that 5.2 does not.

Four generations, one score, four prices

Four GLM generations. All 9/9. Cost goes up, down, up.Same nine executed Python tasks, temperature 0. Bars are ordered by version, not by price.GLM-5 · 9/9$2.30GLM-5.1 · 9/9$4.35 — worstGLM-5.2 · 9/9$1.99 — bestGLM-5.3 · 9/9$3.35 — newestOne scale throughout: 110 px per dollar. GLM-5 and 5.1 have a 204,800-token window; 5.2 and 5.3 have 1,048,576.
The cheapest GLM has been GLM-5.2 through two subsequent releases.

GLM-5.1 was the expensive outlier at $4.35 and 23.6 seconds on 1,327 reasoning tokens. GLM-5.2 cut that to $1.99 and 12.3 seconds. GLM-5.3 gives about half of that saving back. Every one of the four solved all nine tasks, so on this workload the version number tells you nothing about quality and quite a lot about price — in an unhelpful direction.

We keep measuring the same thing across vendors. Grok 4.6 costs 70% more than Grok 4.5 at an identical list price. GPT-5.1-Codex-Max lists below GPT-5.2-Codex and costs 3.09x more. Qwen3.8-Max raised its price 36% and cut the bill anyway. Price per token is half the equation; token volume is the half nobody publishes.

The Ox Alpha question, and our mistake

When the free stealth endpoint stealth/ox-alpha appeared, the most repeated guess was GLM-5.3 Flash. We ran this model — z-ai/glm-5.3, not Flash — on the same suite the same day, found 631 reasoning tokens against Ox Alpha's zero, and treated that as evidence against the GLM guess.

Z.ai confirmed on 2026-08-26 that Ox Alpha was GLM-5.3-Flash. The guess was right and our rebuttal was wrong, for two reasons: Flash is a different model from the one on this page (320B total and 18B active, against 744B and 40B), and reasoning-token behaviour is set by the endpoint rather than the weights. Re-running the model under its own name on the identical tasks produced 1,212 reasoning tokens where the stealth endpoint produced 0. The full correction is here.

Which GLM to run

Against the wider open field GLM is mid-priced: DeepSeek V3.2 scored 9/9 at $0.08 and Qwen3 Coder Next at $0.10. GLM-5.3 ranks 28th cheapest among the 44 models in our set that scored 9 out of 9. If you are running GLM through a coding agent, the GLM Coding Plan page covers the subscription route and using Claude Code with GLM covers the wiring.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived — measured token counts multiplied by the list price on the date shown, not a billing statement. The four generations were measured on different dates at their then-current prices, which is why every figure needs a date. Runs go through OpenRouter. Full method on the methodology page.

What we did not measure

FAQ

Is GLM-5.3 better than GLM-5.2? Not on our nine tasks. Both scored 9 out of 9; 5.3 cost 68% more and was 15% slower.

How much does GLM-5.3 cost? It lists at $1.40 per million input tokens and $4.40 output as of 2026-08-22. On our suite that came to $3.35 per 1,000 tasks.

Which GLM is cheapest per finished task? Of the four we have measured, GLM-5.2 at $1.99.

Was Ox Alpha GLM-5.3? It was GLM-5.3-Flash, a smaller sibling of the model on this page, confirmed by Z.ai on 2026-08-26. We originally argued against the GLM guess and were wrong.

What is GLM-5.3's context window? 1,048,576 tokens, the same as GLM-5.2 and five times GLM-5 and 5.1.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.