Model Reviews

GLM-5.3-Flash Review: A Tenth the Cost of GLM-5.3, and Slower Than All of Them

GLM-5.3-Flash scored 9 out of 9 on our executed Python benchmark at $0.34 per 1,000 tasks. The full-size GLM-5.3 scored the same at $3.35 — so Flash costs a tenth as much for an identical result. It gets there the hard way: at 24.9 seconds it is the slowest of the five GLM generations we have measured, and it emitted 11,884 output tokens against GLM-5.3's 6,640. Nearly twice the tokens, a tenth of the bill, because the list price is $0.075 and $0.25 against $1.40 and $4.40. This is the case where price per token beats volume outright — the opposite of what we have been finding all month. You may also know this model by another name: it shipped anonymously as Ox Alpha first.

Correction, October 3, 2026: single-model benchmark latency is a mean, not a median. Labels were corrected; measured values and the original run dates are unchanged. Measurement details.

Chart of measured cost per 1,000 tasks across five GLM generations with GLM-5.3-Flash far cheapest

Z.ai shipped this model twice: once anonymously in August as a free stealth endpoint, and once under its own name at a promotional price. We measured it both times. This page is about the named version, which is the one you can plan around.

The result

MetricGLM-5.3-Flash
Score9/9
Measured cost / 1,000 tasks$0.34
Mean latency24.9s
Reasoning tokens per call1,212
Output tokens on the suite11,884
List price in / out$0.075 / $0.25 per 1M (promotional)
Context window1,310,720 tokens
Measured2026-08-31

Per task, it ranged from 388 output tokens on flatten to 2,492 on lcs_len, and from 7.9 seconds to 43.3. It is not a consistent model; it is a cheap one that finishes.

Against the rest of the GLM family

Five generations, identical nine tasks, all scoring 9 out of 9:

ModelList price in / outMeasured cost / 1kLatencyReasoning tokensContext
GLM-5.3-Flash$0.075 / $0.25$0.3424.9s1,2121,310,720
GLM-5.2$1.19 / $3.74$1.9912.3s5591,048,576
GLM-5$0.60 / $1.92$2.3016.5s719204,800
GLM-5.3$1.40 / $4.40$3.3514.2s6311,310,720
GLM-5.1$0.97 / $3.04$4.3523.6s1,327204,800
Five GLM models, all 9/9, and one of them costs a tenth of anotherMeasured cost per 1,000 tasks on the identical nine executed Python tasks.GLM-5.3-Flash$0.34GLM-5.2$1.99GLM-5$2.30GLM-5.3$3.35One scale throughout: 135 px per dollar. GLM-5.1 at $4.35 is omitted from the bars for space and appears in the table above.GLM-5.3-Flash is priced promotionally until 2026-09-09; its bar roughly doubles after that.
Every model here solved all nine tasks. The spread is entirely economic.

Cheapest and slowest at the same time

The usual trade is that a cheap model is fast because it thinks less. Flash inverts it. It is the slowest of the five at 24.9 seconds and emits more reasoning than every generation except GLM-5.1. It is cheap anyway, because $0.25 per million output tokens is a seventeenth of GLM-5.3's $4.40.

That is worth pausing on, because it runs against everything else we measured this month. GPT-5.1-Codex-Max lists below its sibling and costs 3.09x more, because volume swamped price. Grok 4.6 matches Grok 4.5's sticker exactly and costs 70% more, same reason. Here the arrow points the other way: Flash emits 1.8x GLM-5.3's output tokens and still bills a tenth, because the price gap is far larger than the volume gap.

Neither factor wins by default. That is the whole argument for measuring cost per finished task instead of reading a price page or a token count in isolation.

The price is promotional until 2026-09-09

The $0.075 and $0.25 we measured against is launch pricing. Z.ai has published that it rises to $0.15 and $0.50 on 2026-09-09, doubling both sides. Our $0.34 per 1,000 tasks will roughly double with it, to somewhere near $0.68 on the same token counts.

Even then it stays the cheapest GLM by a wide margin. But every figure on this page carries a capture date of 2026-08-31 for exactly this reason, and if you are reading this after 2026-09-09 the price you see will not be the price we measured. That is the general case, not a special one — see why every LLM price needs a date attached.

What GLM-5.3-Flash is

Third-party specifications, read 2026-08-31:

Who should run it

Our open-source guide covers the wider field, and the GLM Coding Plan page covers the subscription route if you are running GLM through an agent.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived from measured token counts at the list price captured 2026-08-31, not a billing statement. Runs go through OpenRouter. Full method on the methodology page.

What we did not measure

FAQ

How much does GLM-5.3-Flash cost? $0.075 per million input tokens and $0.25 output at launch pricing, rising to $0.15 and $0.50 on 2026-09-09. On our nine tasks that worked out to $0.34 per 1,000 tasks.

Is GLM-5.3-Flash as good as GLM-5.3? On our tasks, yes — both scored 9 out of 9. Flash cost a tenth as much and took 75% longer.

How fast is GLM-5.3-Flash? 24.9 seconds mean, the slowest of the five GLM models we have measured.

Was GLM-5.3-Flash the Ox Alpha stealth model? Yes, confirmed by Z.ai on 2026-08-26.

What licence is it under? MIT, with weights on Hugging Face.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.