GPT-6 Astra vs Claude Fable 5.1: Identical Price, Identical Score
OpenAI and Anthropic shipped flagships two weeks apart into exactly the same price bracket: $10 per million input tokens and $50 output. We ran both on the same nine executed Python tasks the same day. Both scored 9 out of 9. GPT-6 Astra cost $8.19 per 1,000 tasks; Claude Fable 5.1 cost $8.09. That is a 1.2% difference, which is not a difference. On this workload the two flagships are interchangeable, and anyone telling you otherwise on the basis of a coding benchmark is reading noise. The decisions that do move money are one tier down and one tier up, and both are measurable.
Side by side
| GPT-6 Astra | Claude Fable 5.1 | Difference | |
|---|---|---|---|
| List price in / out | $10 / $50 | $10 / $50 | none |
| Score | 9/9 | 9/9 | none |
| Measured cost / 1,000 tasks | $8.19 | $8.09 | 1.2% |
| Mean latency | 5.6s | 6.8s | 1.2s — Fable takes 21% longer |
| Reasoning tokens per call | 47 | 0 | 47 |
| Input tokens across the suite | 604 | 914 | 51% |
| Output tokens across the suite | 1,353 | 1,273 | 6% |
| Context window | 1,050,000 | 1,000,000 | 5% |
Why the tie is the finding
A 1.2% cost difference on a nine-task suite is below the resolution of the instrument. We run one scored attempt per task and do not average across runs, so a ten-cent gap per thousand tasks is well inside what a re-run could flip. The correct conclusion is that these two models are the same purchase on this workload, not that Anthropic is fractionally cheaper.
That is worth saying plainly because the comparison content you will find elsewhere this month will pick a winner from exactly this kind of margin. If a benchmark cannot separate two models, the honest output is a tie.
Where they actually differ
Three real differences, none of which showed up in the total:
- Astra is 1.2 seconds faster — 5.6 seconds against 6.8, so Fable takes 21% longer. If a person is waiting, that is the tiebreaker.
- Fable 5.1 emits zero reasoning tokens; Astra emits 47. Both are low. At 47 tokens per call the difference is worth fractions of a cent, but the shape matters if you run many turns: a model with no reasoning budget has a bill that tracks only its visible output.
- Astra used 604 input tokens across the suite; Fable used 914 — on identical prompts. That is a 310-token gap from provider-side prompt handling, not from anything we sent. It is small here. It is not always small: Astra Pro injects roughly 1,650 tokens per request, which is how it reaches $35.44 on the same rate card.
The decisions that do move money
If the flagship choice is a coin flip, the money is elsewhere, and both directions are measured.
| Decision | Measured effect |
|---|---|
| Flagship A or flagship B | 1.2% |
| Flagship or the Pro tier above it | 4.3x — $8.19 to $35.44, same score |
| Flagship or Claude Opus 4.8 | 2x — $8.09 to $4.05, same score |
| Flagship or DeepSeek V3.2 | 101x — $8.09 to $0.08, same score |
Every row says 9 out of 9. That is the honest limit of this suite: it establishes a floor, not a ceiling. Nine self-contained Python functions cannot tell you whether a flagship holds a repository in its head better than a small model — and that, not a ten-cent gap, is what you are actually buying at $10 and $50.
So use the measurement for what it is good for: eliminating the differences that are not differences, and pricing the ones that are. Our Astra review and Fable 5.1 review have each model in full.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Both models were run the same day at the same settings, which is what makes a 1.2% comparison meaningful enough to call a tie. Cost is derived from measured token counts at the list price captured 2026-09-15, not a billing statement. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- Long-horizon agentic work, which is the pitch for both models and the one thing this suite cannot see.
- Context windows of 1,050,000 and 1,000,000 tokens. Our prompts are short.
- Anything but Python, and only nine self-contained functions.
- Repeat runs. One scored attempt per task, which is exactly why we call 1.2% a tie rather than a win.
FAQ
Which is better, GPT-6 Astra or Claude Fable 5.1? On our nine executed Python tasks, neither — both scored 9 out of 9 at $8.19 and $8.09 per 1,000 tasks on an identical $10 and $50 rate card.
Is one of them faster? Astra, at 5.6 seconds mean against Fable 5.1's 6.8 — 1.2 seconds, or 21% longer for Fable.
Do they cost the same? Their list prices are identical, and their measured cost differed by 1.2%, which is inside our measurement noise.
What about GPT-6 Astra Pro? Same 9 out of 9, $35.44 per 1,000 tasks — 4.3x, because it prepends about 1,650 tokens of hidden prompt per request.
Should I pay for either? Not for work like our suite. Claude Opus 4.8 scored the same at half the cost, and DeepSeek V3.2 at a hundredth.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab