Claude Fable 5.1 Review: Nine of Nine, and No Filter This Time
Claude Fable 5.1 scored 9 out of 9 on our executed Python benchmark at $8.09 per 1,000 tasks, a 6.8-second mean, and zero reasoning tokens on every task. That last number is the shape of the model: it answers without a visible thinking budget and still clears the suite. The result also closes a loop for us. Claude Fable 5 never finished this benchmark — a content filter returned empty bodies on four of nine tasks, which our harness originally scored as wrong answers before we fixed it. 5.1 ran clean on the first attempt, nine for nine, no retries.
Most point releases are hard to write about because nothing measurable changes. This one has a clean before-and-after, because the previous version did not complete our suite at all.
The result
| Metric | Claude Fable 5.1 |
|---|---|
| Score | 9/9 |
| Measured cost / 1,000 tasks | $8.09 |
| Mean latency | 6.8s |
| Reasoning tokens per call | 0 |
| Tokens across the suite | 914 in / 1,273 out |
| List price in / out | $10 / $50 per 1M |
| Context window | 1,000,000 |
| Measured | 2026-09-16 |
Among the 54 models in our set that scored 9 out of 9, it ranks 47th cheapest and 21st fastest. That is where a $10-and-$50 model belongs on a suite this easy; the interesting parts are elsewhere.
What happened with Fable 5
We ran Claude Fable 5 on 2026-07-30. It is still marked excluded in our benchmark data, and the reason matters.
| Claude Fable 5 | Claude Fable 5.1 | |
|---|---|---|
| Tasks that returned an answer | 5 of 9 | 9 of 9 |
| Of those, passed | 4 | 9 |
| Recorded score | 4/5 — excluded | 9/9 |
| Reasoning tokens | 0 | 0 |
| Tokens across the suite | 438 in / 539 out | 914 in / 1,273 out |
| List price | $10 / $50 | $10 / $50 |
Four of the nine tasks came back with an empty body and a content-filter finish reason. Our harness at the time scored a non-answer as a wrong answer, which would have published a headline of 4 out of 9 for Anthropic's flagship — a number about our instrument, not the model. We caught it, fixed the harness to separate an API-layer failure from a wrong answer, and wrote the whole thing up in the content filter post.
Fable 5.1 triggered none of it. Nine tasks, nine answers, no retries, no empty bodies. We are not claiming Anthropic changed a filter in response to anything — we have no visibility into that, and these are two runs six weeks apart on prompts about parsing CSV lines and merging intervals. What we can say is narrow and checkable: the behaviour that broke the earlier run did not recur.
Zero reasoning tokens at a frontier price
Fable 5.1 reported 0 reasoning tokens across all nine tasks. Of the models in our set, that puts it in a small group — and it is the only one in that group carrying a $10-and-$50 rate card.
That shapes the economics in a way the price page hides. Reasoning tokens bill at the output rate, so a model that emits none has a bill that tracks its visible answer and nothing else. On our suite Fable 5.1 produced 1,273 output tokens across nine tasks. Gemini 3.8 Flash produced 9,168 for the same nine answers, at a seventh of the per-token price — and still landed at less than half the cost. Expensive and terse versus cheap and verbose is a real trade, and which side wins depends entirely on your prompt shape.
We have been measuring that trade all summer, in both directions: GPT-5.1-Codex-Max lists below its sibling and costs 3.09x more on volume alone, while GLM-5.3-Flash emits 1.8x the tokens of its full-size sibling and bills a tenth.
Against Astra and Opus 5
| Model | Score | Measured cost / 1k | Latency | Reasoning tokens | List price |
|---|---|---|---|---|---|
| Claude Opus 4.8 | 9/9 | $4.05 | 6.1s | 0 | — |
| Claude Opus 5 | 9/9 | $5.64 | 5.3s | 6 | $5 / $25 |
| Claude Fable 5.1 | 9/9 | $8.09 | 6.8s | 0 | $10 / $50 |
| GPT-6 Astra | 9/9 | $8.19 | 5.6s | 47 | $10 / $50 |
Two readings. Within Anthropic's own lineup, Fable 5.1 costs twice what Opus 4.8 costs for the same 9 out of 9 and is slower than Opus 5 — on this suite, the flagship is the worst value of the three. Against OpenAI, Fable 5.1 and GPT-6 Astra land ten cents apart per thousand tasks on an identical rate card, which makes them the same purchase on this workload. We take that pair apart in the head-to-head.
Who should pay $10 and $50
- Not for work like our suite. Claude Haiku 4.5 scored the same 9 out of 9 at $0.94, and DeepSeek V3.2 at $0.08. Nine self-contained functions do not separate a flagship from a competent small model.
- The case for it is the work we cannot measure — long multi-file sessions, holding a large repository in context, recovering from its own mistakes. Nothing on this page speaks to that, in either direction.
- If you are on Fable 5, the upgrade is straightforward: same price, and it completes work the previous version returned nothing for.
- If your prompts are short and your outputs long, the zero-reasoning behaviour is worth real money relative to a chatty cheaper model. Measure your own shape.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure — an empty body, a filter, a timeout — is recorded separately from a wrong answer, which is the fix Fable 5 forced on us. Cost is derived from measured token counts at the list price captured 2026-09-15, not a billing statement. Runs go through OpenRouter. Full method on the methodology page, and prices move.
What we did not measure
- The 1,000,000-token context. Our prompts are short.
- Long agentic sessions, which is where a flagship earns its price and our suite is silent.
- Why Fable 5 filtered and 5.1 did not. We observed two runs six weeks apart. We have no view into the filter itself and make no claim about what changed.
- Anything but Python, and only nine self-contained functions.
- Repeat runs. One scored attempt per task. Single-run figures, not averages.
FAQ
How much does Claude Fable 5.1 cost? $10 per million input tokens and $50 output as of 2026-09-15. On our nine tasks that came to $8.09 per 1,000 tasks.
Is Fable 5.1 better than Fable 5? On our suite, measurably: 5.1 answered all nine tasks where 5 returned empty bodies on four of them.
How many reasoning tokens does Fable 5.1 use? Zero, across all nine tasks.
Fable 5.1 or GPT-6 Astra? On this workload they are the same purchase — both 9 out of 9, $8.09 against $8.19, on an identical $10 and $50 rate card.
Is it better value than Claude Opus 4.8? Not on these tasks. Opus 4.8 scored the same 9 out of 9 at $4.05.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab