GPT-6 Astra Pro Review: The Same Score, at 4.3x the Cost
GPT-6 Astra Pro consumed 17,011 input tokens to answer nine short Python prompts. Plain GPT-6 Astra answered the identical nine on 604. Both models scored 9 out of 9, both list at $10 in and $50 out per million, and the derived cost came out at $35.44 per 1,000 tasks for Pro against $8.19 for Astra — 4.3x, from a rate card that is identical to the cent. Nothing in the Pro tier thought much harder: 111 reasoning tokens per call against 47. The bulk of the multiple is input you did not write.
A price page is supposed to be the thing you can plan against. Two endpoints from one vendor, same numbers on the card, same score on the suite, and one of them costs four times the other. The card was never wrong. It was just never the whole input.
The result
| Metric | GPT-6 Astra Pro | GPT-6 Astra |
|---|---|---|
| Score | 9/9 | 9/9 |
| Derived cost / 1,000 tasks | $35.44 | $8.19 |
| Mean latency | 8.4s | 5.6s |
| Reasoning tokens per call | 111 | 47 |
| Input tokens across the suite | 17,011 | 604 |
| Output tokens across the suite | 2,977 | 1,353 |
| List price in / out per 1M | $10 / $50 | $10 / $50 |
| Context window | 1,050,000 | 1,050,000 |
| Measured | 2026-09-16 | 2026-09-16 |
Both prices are derived from measured token counts at the list price captured 2026-09-15. Among the 56 models in our set that score 9 out of 9, Astra Pro ranks 55th cheapest and 29th fastest. There is exactly one 9/9 model in our data that costs more.
The cost multiple is $35.44 ÷ $8.19 ≈ 4.33, and the input multiple is 17,011 ÷ 604 ≈ 28.2. Those two numbers are the whole review. Everything below is why the second one produces the first.
It is an input bill, not a thinking bill
The intuitive story for an expensive premium tier is that it thinks longer — more reasoning tokens, more output, a bigger completion to pay for. That story is measurable, and on our suite it does not get you to 4.3x. Astra Pro emitted 111 reasoning tokens per call and 2,977 output tokens across the whole suite. For scale, Ling 3.0 Flash VL emitted 3,299 output tokens on the same nine tasks and cost $0.07 per 1,000. Pro's output did grow against plain Astra — 2,977 against 1,353, a factor of 2.2. Its input grew by a factor of 28.2. Both sides of the bill moved; only one moved by an order of magnitude.
Put a third model on the same card and the picture sharpens. Claude Fable 5.1 also lists at $10 in and $50 out, also scored 9 out of 9, and spent 914 input tokens across the suite. Astra at 604 and Fable at 914 are both in the range you would expect from nine short function specs. Astra Pro is not in that range at all.
What one request looks like
Suite totals could in principle hide something mundane — a retry, a longer task template, a stray system message on our side. So we sent one single benchmark prompt to each endpoint and read the raw usage payload back. Those figures are published in full in the GPT-6 Astra review; here is the row that matters:
| Model | Message we sent | Prompt tokens billed | of which cached |
|---|---|---|---|
| GPT-6 Astra | 65 tokens | 65 | 0 |
| GPT-6 Astra Pro | 65 tokens | 1,721 | 1,389 |
Astra billed exactly what we sent. Astra Pro billed 1,721 for the same 65-token message — roughly 26 times the input we actually wrote, as that article works through. 1,389 of those tokens were cached, which is the part most people misread. Cached input is discounted input on most providers, not free input, and it still shows up in the usage object you are billed against. The cache tells you something else too: the injected block is stable across calls, which is what a fixed provider-side system prompt looks like from outside.
We measured the size of that block. We did not read its text, and we are not going to pretend we know what is in it.
The company it keeps at the top of the price list
Astra Pro is the second most expensive model in our set that clears the whole suite. The most expensive is sakana/fugu-ultra-v2 at $57.06 per 1,000 tasks, and it got there the same way:
| Model | Score | Derived cost / 1k | List price in / out | Input tokens | Output tokens |
|---|---|---|---|---|---|
| sakana/fugu-ultra-v2 | 9/9 | $57.06 | $5 / $30 | 60,645 | 7,011 |
| GPT-6 Astra Pro | 9/9 | $35.44 | $10 / $50 | 17,011 | 2,977 |
| Claude Fable 5.1 | 9/9 | $8.09 | $10 / $50 | 914 | 1,273 |
| GPT-6 Astra | 9/9 | $8.19 | $10 / $50 | 604 | 1,353 |
Read the list-price column against the cost column. Fugu Ultra v2 lists at half Astra Pro's input rate, and $30 against $50 on output, and still costs more per 1,000 tasks, because it carried 60,645 input tokens through nine short prompts. Both of the two priciest 9/9 endpoints we have measured reached the top of the cost table through input volume rather than through their rate cards. If you are shopping by rate card at the premium end, the rate card is the least predictive column on the page.
This is also the honest read on watching list prices move: a vendor can leave the card untouched for months while the effective cost of a request changes, because the card prices tokens and the endpoint decides how many tokens a request is.
The alias points at the cheap one
One small thing worth knowing if you route through a floating tag. In the alias map we captured on 2026-09-15 and published in the piece on what latest actually resolves to, ~openai/gpt-astra-latest resolves to openai/gpt-6-astra — the $8.19 endpoint, not the $35.44 one. That is the safe default in this particular family, and it is the opposite of the usual worry about floating tags silently upgrading you into a more expensive tier. Do not generalise it; check the map for whichever family you route through.
When Pro is still the right buy
Here is the part we are obliged to say plainly. Nine self-contained Python functions cannot separate a frontier model from a competent small one. Fifty-six models in our set clear this suite, including GPT-5.4 mini at $0.53 and DeepSeek V3.2 at $0.08. A benchmark that an eight-cent model aces is not a benchmark that can tell you a Pro tier is worthless. It genuinely cannot.
What it can tell you is narrower and still useful:
- The overhead is real and it is fixed per call. 1,721 billed prompt tokens for a 65-token message is not a rounding error on short requests.
- It amortises with prompt length. On a 20,000-token request the same injected block is noise. On a short classification or extraction call fired thousands of times a day, it is most of your bill.
- It bought nothing measurable here. Identical score, and 8.4s against 5.6s mean — Pro was slower on this workload, not faster.
So the decision rule is not “never buy Pro”. It is: if your prompts are short and repetitive, price the overhead before you commit, because your rate card will not price it for you. If your workload is the long-horizon agentic work the tier is actually built for, our suite has no opinion — measure it yourself on tasks that stress it. For the other end of the market entirely, our cheap coding roundup has the field.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived — measured input and output token counts multiplied by the list price captured 2026-09-15 — not a billing statement, and your invoice will differ with caching discounts and negotiated rates. The single-request prompt-token figures come from a separate probe whose raw usage payload we read directly. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- The thing Pro is sold for. Long-horizon, multi-step, tool-using agent work. Our nine tasks are single-turn and self-contained. This review measures the cost floor of a tier, not its ceiling.
- The contents of the injected prompt. We have its size and the fact that 1,389 of 1,721 tokens were cached. We have not read a word of it, and we are not inferring its purpose.
- Whether the cached portion is discounted on your account. Our derived cost applies the standard $10 input rate to every input token, cached or not. If your provider discounts cache reads, your real bill is lower than $35.44 — and still well above $8.19.
- The 1,050,000-token context. Our prompts are a few hundred tokens. We have tested none of it.
- Repeat runs. One scored attempt per task, one run per model. Means across tasks from a single run, not averages across repeated runs, so the 8.4s mean latency should be read loosely.
- Whether the overhead is stable over time. We probed once, on one day. A provider-side prompt can be rewritten between deploys without any announcement, which would move every number in this article.
FAQ
How much does GPT-6 Astra Pro cost? $10 per million input tokens and $50 output as of 2026-09-15 — identical to plain Astra. On our nine tasks that worked out to $35.44 per 1,000 tasks, against $8.19 for Astra.
Why is GPT-6 Astra Pro more expensive at the same list price? Input volume. It used 17,011 input tokens across nine prompts where Astra used 604, because a large provider-side prompt is prepended to every request.
Is GPT-6 Astra Pro better than GPT-6 Astra? Not on our suite. Both scored 9 out of 9, and Pro was slower at 8.4s mean against 5.6s. Nine short Python functions cannot test what the Pro tier is built for, so this is a cost finding, not a capability verdict.
Does caching make it cheap again? It helps. 1,389 of 1,721 prompt tokens on our probe were cached. Cached input is discounted, not free, and our derived figure does not apply any cache discount.
Which endpoint does the latest alias give me? ~openai/gpt-astra-latest resolved to openai/gpt-6-astra when we captured the map on 2026-09-15 — the cheaper one.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab