Model Reviews

GPT-6 Astra Review: The Name Landed, and So Did a Hidden Prompt

GPT-6 Astra is out, and it is called GPT-6 — the model ships as openai/gpt-6-astra, settling a naming question OpenAI had explicitly left open two weeks ago. On our executed Python benchmark it scored 9 out of 9 at $8.19 per 1,000 tasks in 5.6 seconds, emitting only 47 reasoning tokens per call. The finding worth your attention is the Pro tier. GPT-6 Astra Pro scored the identical 9 out of 9 and cost $35.44 — 4.3x — on the same list price. It is not thinking harder. It is prepending roughly 1,650 tokens of hidden prompt to every request you send, which we confirmed by sending a 65-token message and reading the usage back.

DataLLM Lab article cover: GPT-6 Astra Review: The Name Landed, and So Did a Hidden Prompt

Two weeks ago Astra was an announced model with no name, no price and no endpoint. It now has all three, so the interesting question moves from what it might be to what it costs to use.

The result

MetricGPT-6 AstraGPT-6 Astra Pro
Score9/99/9
Measured cost / 1,000 tasks$8.19$35.44
Mean latency5.6s8.4s
Reasoning tokens per call47111
Input tokens across the suite60417,011
Output tokens across the suite1,3532,977
List price in / out$10 / $50$10 / $50
Context window1,050,0001,050,000

Astra itself is a well-behaved model on this suite: fast, terse, and it does not over-think simple problems — 47 reasoning tokens per call is low for a system marketed on long-horizon reasoning. Among the 54 models in our set that scored 9 out of 9 it ranks 48th cheapest and 15th fastest. It is a premium model that behaves like one.

Now read the input-token row again.

The Pro tier and the hidden prompt

Astra consumed 604 input tokens across nine prompts. Astra Pro consumed 17,011 on the identical nine. That is 28x the input for the same questions, and it is the entire reason the bill is 4.3x.

A 28x gap on identical text has to come from somewhere, so we sent both models a single benchmark prompt and read the raw usage:

ModelPrompt tokens billedof which cachedCompletion
GPT-6 Astra65058
GPT-6 Astra Pro1,7211,389131

The message we sent was 65 tokens. Astra billed 65. Astra Pro billed 1,721 — roughly 1,656 tokens of prompt you did not write, injected before your request on every call. Most of it is cached, which makes it cheaper on repeat, but cached input is still billed input.

You sent 65 tokens. The Pro tier billed you for 1,721.Same prompt, same moment. The difference is prompt the provider adds before yours.PROMPT TOKENS BILLED, ONE REQUESTGPT-6 Astra65 — exactly what we sentGPT-6 Astra Pro1,389 cached1,721 totalMEASURED COST PER 1,000 TASKS · BOTH 9/9GPT-6 Astra$8.19GPT-6 Astra Pro$35.44
The scaffolding is the product. It is also the bill.

This is not an accusation of anything hidden in a sinister sense — a “Pro” agentic tier plausibly needs a large system prompt to do what it does, and OpenAI is entitled to build it that way. It is a statement about what a price page cannot tell you. Both tiers list at $10 and $50 per million tokens. Nothing on that page says one of them starts every conversation 1,650 tokens in the red.

If your prompts are short, the overhead dominates. On a 65-token request, Astra Pro's injected prompt is 26 times your actual input. On a 20,000-token request it is noise. Size your expectations by your own prompt length, not by the rate card.

Against Claude Fable 5.1, at the same price

Anthropic shipped Claude Fable 5.1 into the same price bracket — $10 and $50 per million, identical to Astra. Same suite, same day:

ModelScoreMeasured cost / 1kLatencyReasoning tokens
GPT-6 Astra9/9$8.195.6s47
Claude Fable 5.19/9$8.096.8s0
GPT-6 Astra Pro9/9$35.448.4s111

Astra and Fable 5.1 finish within ten cents per thousand tasks of each other on an identical rate card. On this workload they are the same purchase. We take that pair apart properly in the head-to-head.

What happened to the naming question

On 2026-08-31 we wrote that Astra was confirmed but its name was not — reporting had OpenAI undecided between GPT-6, a GPT-5 point release, and a fourth class beside Sol, Terra and Luna. The endpoint settles it: the model ships as gpt-6-astra, so it is GPT-6, and there is a Pro variant alongside it. The ~openai/gpt-astra-latest alias resolves to openai/gpt-6-astra, which is the same answer from a second direction.

Worth noting for anyone who read the rumour coverage: the widely repeated ultima-alpha codename and the 3–10 September window never had a named source, and the model arrived outside that window.

Is it worth $8.19?

On nine self-contained Python functions, plainly no — DeepSeek V3.2 scored the same 9 out of 9 at $0.08, which is 102x cheaper, and GPT-5.4 mini did it at $0.53 in 2.3 seconds. Our suite cannot distinguish a frontier model from a competent small one, and we would not pretend otherwise.

What our suite can tell you is the floor price of the capability you are buying, and whether the tier above is buying you anything measurable. On that second question the answer is concrete: Astra Pro cost 4.3x and scored identically. If you are considering it, the burden of proof is on your workload — measure it on tasks that actually stress long-horizon reasoning before paying the multiple.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived — measured input and output token counts multiplied by the list price captured 2026-09-15, not a billing statement. The hidden-prompt figures come from a separate single request whose raw usage payload we read directly. Runs go through OpenRouter. Full method on the methodology page, and prices move.

What we did not measure

FAQ

Is GPT-6 out? Yes. It ships as GPT-6 Astra, with a GPT-6 Astra Pro variant alongside it.

How much does GPT-6 Astra cost? $10 per million input tokens and $50 output as of 2026-09-15. On our nine tasks that worked out to $8.19 per 1,000 tasks.

Is GPT-6 Astra Pro worth the extra? Not on our benchmark. It scored the same 9 out of 9 and cost $35.44 against $8.19.

Why is Astra Pro so much more expensive at the same list price? It prepends roughly 1,650 tokens of provider-side prompt to every request. We sent a 65-token message and were billed for 1,721.

GPT-6 Astra or Claude Fable 5.1? On this suite they are indistinguishable — both 9 out of 9, $8.19 against $8.09, on the same $10 and $50 rate card.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.