GPT-6 Astra Review: The Name Landed, and So Did a Hidden Prompt
GPT-6 Astra is out, and it is called GPT-6 — the model ships as openai/gpt-6-astra, settling a naming question OpenAI had explicitly left open two weeks ago. On our executed Python benchmark it scored 9 out of 9 at $8.19 per 1,000 tasks in 5.6 seconds, emitting only 47 reasoning tokens per call. The finding worth your attention is the Pro tier. GPT-6 Astra Pro scored the identical 9 out of 9 and cost $35.44 — 4.3x — on the same list price. It is not thinking harder. It is prepending roughly 1,650 tokens of hidden prompt to every request you send, which we confirmed by sending a 65-token message and reading the usage back.
Two weeks ago Astra was an announced model with no name, no price and no endpoint. It now has all three, so the interesting question moves from what it might be to what it costs to use.
The result
| Metric | GPT-6 Astra | GPT-6 Astra Pro |
|---|---|---|
| Score | 9/9 | 9/9 |
| Measured cost / 1,000 tasks | $8.19 | $35.44 |
| Mean latency | 5.6s | 8.4s |
| Reasoning tokens per call | 47 | 111 |
| Input tokens across the suite | 604 | 17,011 |
| Output tokens across the suite | 1,353 | 2,977 |
| List price in / out | $10 / $50 | $10 / $50 |
| Context window | 1,050,000 | 1,050,000 |
Astra itself is a well-behaved model on this suite: fast, terse, and it does not over-think simple problems — 47 reasoning tokens per call is low for a system marketed on long-horizon reasoning. Among the 54 models in our set that scored 9 out of 9 it ranks 48th cheapest and 15th fastest. It is a premium model that behaves like one.
Now read the input-token row again.
The Pro tier and the hidden prompt
Astra consumed 604 input tokens across nine prompts. Astra Pro consumed 17,011 on the identical nine. That is 28x the input for the same questions, and it is the entire reason the bill is 4.3x.
A 28x gap on identical text has to come from somewhere, so we sent both models a single benchmark prompt and read the raw usage:
| Model | Prompt tokens billed | of which cached | Completion |
|---|---|---|---|
| GPT-6 Astra | 65 | 0 | 58 |
| GPT-6 Astra Pro | 1,721 | 1,389 | 131 |
The message we sent was 65 tokens. Astra billed 65. Astra Pro billed 1,721 — roughly 1,656 tokens of prompt you did not write, injected before your request on every call. Most of it is cached, which makes it cheaper on repeat, but cached input is still billed input.
This is not an accusation of anything hidden in a sinister sense — a “Pro” agentic tier plausibly needs a large system prompt to do what it does, and OpenAI is entitled to build it that way. It is a statement about what a price page cannot tell you. Both tiers list at $10 and $50 per million tokens. Nothing on that page says one of them starts every conversation 1,650 tokens in the red.
If your prompts are short, the overhead dominates. On a 65-token request, Astra Pro's injected prompt is 26 times your actual input. On a 20,000-token request it is noise. Size your expectations by your own prompt length, not by the rate card.
Against Claude Fable 5.1, at the same price
Anthropic shipped Claude Fable 5.1 into the same price bracket — $10 and $50 per million, identical to Astra. Same suite, same day:
| Model | Score | Measured cost / 1k | Latency | Reasoning tokens |
|---|---|---|---|---|
| GPT-6 Astra | 9/9 | $8.19 | 5.6s | 47 |
| Claude Fable 5.1 | 9/9 | $8.09 | 6.8s | 0 |
| GPT-6 Astra Pro | 9/9 | $35.44 | 8.4s | 111 |
Astra and Fable 5.1 finish within ten cents per thousand tasks of each other on an identical rate card. On this workload they are the same purchase. We take that pair apart properly in the head-to-head.
What happened to the naming question
On 2026-08-31 we wrote that Astra was confirmed but its name was not — reporting had OpenAI undecided between GPT-6, a GPT-5 point release, and a fourth class beside Sol, Terra and Luna. The endpoint settles it: the model ships as gpt-6-astra, so it is GPT-6, and there is a Pro variant alongside it. The ~openai/gpt-astra-latest alias resolves to openai/gpt-6-astra, which is the same answer from a second direction.
Worth noting for anyone who read the rumour coverage: the widely repeated ultima-alpha codename and the 3–10 September window never had a named source, and the model arrived outside that window.
Is it worth $8.19?
On nine self-contained Python functions, plainly no — DeepSeek V3.2 scored the same 9 out of 9 at $0.08, which is 102x cheaper, and GPT-5.4 mini did it at $0.53 in 2.3 seconds. Our suite cannot distinguish a frontier model from a competent small one, and we would not pretend otherwise.
What our suite can tell you is the floor price of the capability you are buying, and whether the tier above is buying you anything measurable. On that second question the answer is concrete: Astra Pro cost 4.3x and scored identically. If you are considering it, the burden of proof is on your workload — measure it on tasks that actually stress long-horizon reasoning before paying the multiple.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived — measured input and output token counts multiplied by the list price captured 2026-09-15, not a billing statement. The hidden-prompt figures come from a separate single request whose raw usage payload we read directly. Runs go through OpenRouter. Full method on the methodology page, and prices move.
What we did not measure
- Long-horizon multi-agent work, which is the entire pitch. Nine single-turn functions cannot test a model built to run for hours.
- The 1,050,000-token context. Our prompts are short.
- What the injected prompt contains. We measured its size, not its text.
- Whether the Pro overhead pays off elsewhere. On this suite it bought nothing; on a harder workload it might be the whole point.
- Repeat runs. One scored attempt per task. Single-run figures, not averages.
FAQ
Is GPT-6 out? Yes. It ships as GPT-6 Astra, with a GPT-6 Astra Pro variant alongside it.
How much does GPT-6 Astra cost? $10 per million input tokens and $50 output as of 2026-09-15. On our nine tasks that worked out to $8.19 per 1,000 tasks.
Is GPT-6 Astra Pro worth the extra? Not on our benchmark. It scored the same 9 out of 9 and cost $35.44 against $8.19.
Why is Astra Pro so much more expensive at the same list price? It prepends roughly 1,650 tokens of provider-side prompt to every request. We sent a 65-token message and were billed for 1,721.
GPT-6 Astra or Claude Fable 5.1? On this suite they are indistinguishable — both 9 out of 9, $8.19 against $8.09, on the same $10 and $50 rate card.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab