Model Reviews

GPT-6 Astra Pro Review: The Same Score, at 4.3x the Cost

GPT-6 Astra Pro consumed 17,011 input tokens to answer nine short Python prompts. Plain GPT-6 Astra answered the identical nine on 604. Both models scored 9 out of 9, both list at $10 in and $50 out per million, and the derived cost came out at $35.44 per 1,000 tasks for Pro against $8.19 for Astra — 4.3x, from a rate card that is identical to the cent. Nothing in the Pro tier thought much harder: 111 reasoning tokens per call against 47. The bulk of the multiple is input you did not write.

DataLLM Lab article cover: GPT-6 Astra Pro Review: The Same Score, at 4.3x the Cost

A price page is supposed to be the thing you can plan against. Two endpoints from one vendor, same numbers on the card, same score on the suite, and one of them costs four times the other. The card was never wrong. It was just never the whole input.

The result

MetricGPT-6 Astra ProGPT-6 Astra
Score9/99/9
Derived cost / 1,000 tasks$35.44$8.19
Mean latency8.4s5.6s
Reasoning tokens per call11147
Input tokens across the suite17,011604
Output tokens across the suite2,9771,353
List price in / out per 1M$10 / $50$10 / $50
Context window1,050,0001,050,000
Measured2026-09-162026-09-16

Both prices are derived from measured token counts at the list price captured 2026-09-15. Among the 56 models in our set that score 9 out of 9, Astra Pro ranks 55th cheapest and 29th fastest. There is exactly one 9/9 model in our data that costs more.

The cost multiple is $35.44 ÷ $8.19 ≈ 4.33, and the input multiple is 17,011 ÷ 604 ≈ 28.2. Those two numbers are the whole review. Everything below is why the second one produces the first.

It is an input bill, not a thinking bill

The intuitive story for an expensive premium tier is that it thinks longer — more reasoning tokens, more output, a bigger completion to pay for. That story is measurable, and on our suite it does not get you to 4.3x. Astra Pro emitted 111 reasoning tokens per call and 2,977 output tokens across the whole suite. For scale, Ling 3.0 Flash VL emitted 3,299 output tokens on the same nine tasks and cost $0.07 per 1,000. Pro's output did grow against plain Astra — 2,977 against 1,353, a factor of 2.2. Its input grew by a factor of 28.2. Both sides of the bill moved; only one moved by an order of magnitude.

Same nine prompts. Same $10 / $50 rate card. Same 9/9.Input tokens consumed across the suite. The bill follows this chart, not the price page.GPT-6 Astra604Claude Fable 5.1914GPT-6 Astra Pro17,011One scale throughout: 0.03 px per input token — 604 draws 18.12 px, 914 draws 27.42 px, 17,011 draws 510.33 px. Prices and token counts captured 2026-09-15 and 2026-09-16.
Three endpoints on one rate card. Only one of them starts every call already loaded.

Put a third model on the same card and the picture sharpens. Claude Fable 5.1 also lists at $10 in and $50 out, also scored 9 out of 9, and spent 914 input tokens across the suite. Astra at 604 and Fable at 914 are both in the range you would expect from nine short function specs. Astra Pro is not in that range at all.

What one request looks like

Suite totals could in principle hide something mundane — a retry, a longer task template, a stray system message on our side. So we sent one single benchmark prompt to each endpoint and read the raw usage payload back. Those figures are published in full in the GPT-6 Astra review; here is the row that matters:

ModelMessage we sentPrompt tokens billedof which cached
GPT-6 Astra65 tokens650
GPT-6 Astra Pro65 tokens1,7211,389

Astra billed exactly what we sent. Astra Pro billed 1,721 for the same 65-token message — roughly 26 times the input we actually wrote, as that article works through. 1,389 of those tokens were cached, which is the part most people misread. Cached input is discounted input on most providers, not free input, and it still shows up in the usage object you are billed against. The cache tells you something else too: the injected block is stable across calls, which is what a fixed provider-side system prompt looks like from outside.

We measured the size of that block. We did not read its text, and we are not going to pretend we know what is in it.

The company it keeps at the top of the price list

Astra Pro is the second most expensive model in our set that clears the whole suite. The most expensive is sakana/fugu-ultra-v2 at $57.06 per 1,000 tasks, and it got there the same way:

ModelScoreDerived cost / 1kList price in / outInput tokensOutput tokens
sakana/fugu-ultra-v29/9$57.06$5 / $3060,6457,011
GPT-6 Astra Pro9/9$35.44$10 / $5017,0112,977
Claude Fable 5.19/9$8.09$10 / $509141,273
GPT-6 Astra9/9$8.19$10 / $506041,353

Read the list-price column against the cost column. Fugu Ultra v2 lists at half Astra Pro's input rate, and $30 against $50 on output, and still costs more per 1,000 tasks, because it carried 60,645 input tokens through nine short prompts. Both of the two priciest 9/9 endpoints we have measured reached the top of the cost table through input volume rather than through their rate cards. If you are shopping by rate card at the premium end, the rate card is the least predictive column on the page.

This is also the honest read on watching list prices move: a vendor can leave the card untouched for months while the effective cost of a request changes, because the card prices tokens and the endpoint decides how many tokens a request is.

The alias points at the cheap one

One small thing worth knowing if you route through a floating tag. In the alias map we captured on 2026-09-15 and published in the piece on what latest actually resolves to, ~openai/gpt-astra-latest resolves to openai/gpt-6-astra — the $8.19 endpoint, not the $35.44 one. That is the safe default in this particular family, and it is the opposite of the usual worry about floating tags silently upgrading you into a more expensive tier. Do not generalise it; check the map for whichever family you route through.

When Pro is still the right buy

Here is the part we are obliged to say plainly. Nine self-contained Python functions cannot separate a frontier model from a competent small one. Fifty-six models in our set clear this suite, including GPT-5.4 mini at $0.53 and DeepSeek V3.2 at $0.08. A benchmark that an eight-cent model aces is not a benchmark that can tell you a Pro tier is worthless. It genuinely cannot.

What it can tell you is narrower and still useful:

So the decision rule is not “never buy Pro”. It is: if your prompts are short and repetitive, price the overhead before you commit, because your rate card will not price it for you. If your workload is the long-horizon agentic work the tier is actually built for, our suite has no opinion — measure it yourself on tasks that stress it. For the other end of the market entirely, our cheap coding roundup has the field.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived — measured input and output token counts multiplied by the list price captured 2026-09-15 — not a billing statement, and your invoice will differ with caching discounts and negotiated rates. The single-request prompt-token figures come from a separate probe whose raw usage payload we read directly. Runs go through OpenRouter. Full method on the methodology page.

What we did not measure

FAQ

How much does GPT-6 Astra Pro cost? $10 per million input tokens and $50 output as of 2026-09-15 — identical to plain Astra. On our nine tasks that worked out to $35.44 per 1,000 tasks, against $8.19 for Astra.

Why is GPT-6 Astra Pro more expensive at the same list price? Input volume. It used 17,011 input tokens across nine prompts where Astra used 604, because a large provider-side prompt is prepended to every request.

Is GPT-6 Astra Pro better than GPT-6 Astra? Not on our suite. Both scored 9 out of 9, and Pro was slower at 8.4s mean against 5.6s. Nine short Python functions cannot test what the Pro tier is built for, so this is a cost finding, not a capability verdict.

Does caching make it cheap again? It helps. 1,389 of 1,721 prompt tokens on our probe were cached. Cached input is discounted, not free, and our derived figure does not apply any cache discount.

Which endpoint does the latest alias give me? ~openai/gpt-astra-latest resolved to openai/gpt-6-astra when we captured the map on 2026-09-15 — the cheaper one.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.