Model Reviews

Claude Fable 5.1 Review: Nine of Nine, and No Filter This Time

Claude Fable 5.1 scored 9 out of 9 on our executed Python benchmark at $8.09 per 1,000 tasks, a 6.8-second mean, and zero reasoning tokens on every task. That last number is the shape of the model: it answers without a visible thinking budget and still clears the suite. The result also closes a loop for us. Claude Fable 5 never finished this benchmark — a content filter returned empty bodies on four of nine tasks, which our harness originally scored as wrong answers before we fixed it. 5.1 ran clean on the first attempt, nine for nine, no retries.

DataLLM Lab article cover: Claude Fable 5.1 Review: Nine of Nine, and No Filter This Time

Most point releases are hard to write about because nothing measurable changes. This one has a clean before-and-after, because the previous version did not complete our suite at all.

The result

MetricClaude Fable 5.1
Score9/9
Measured cost / 1,000 tasks$8.09
Mean latency6.8s
Reasoning tokens per call0
Tokens across the suite914 in / 1,273 out
List price in / out$10 / $50 per 1M
Context window1,000,000
Measured2026-09-16

Among the 54 models in our set that scored 9 out of 9, it ranks 47th cheapest and 21st fastest. That is where a $10-and-$50 model belongs on a suite this easy; the interesting parts are elsewhere.

What happened with Fable 5

We ran Claude Fable 5 on 2026-07-30. It is still marked excluded in our benchmark data, and the reason matters.

Claude Fable 5Claude Fable 5.1
Tasks that returned an answer5 of 99 of 9
Of those, passed49
Recorded score4/5 — excluded9/9
Reasoning tokens00
Tokens across the suite438 in / 539 out914 in / 1,273 out
List price$10 / $50$10 / $50
Nine tasks, six weeks apart, same rate cardEach square is one task. Grey = no answer returned at all, not a wrong answer.Claude Fable 52026-07-304 passed · 1 wrong · 4 emptyClaude Fable 5.12026-09-169 passedPale squares are tasks that returned an empty body with a content-filter finish reason. Our harness originally scored those as wrong answers.
The four pale squares are the reason Fable 5 carries no score in our data.

Four of the nine tasks came back with an empty body and a content-filter finish reason. Our harness at the time scored a non-answer as a wrong answer, which would have published a headline of 4 out of 9 for Anthropic's flagship — a number about our instrument, not the model. We caught it, fixed the harness to separate an API-layer failure from a wrong answer, and wrote the whole thing up in the content filter post.

Fable 5.1 triggered none of it. Nine tasks, nine answers, no retries, no empty bodies. We are not claiming Anthropic changed a filter in response to anything — we have no visibility into that, and these are two runs six weeks apart on prompts about parsing CSV lines and merging intervals. What we can say is narrow and checkable: the behaviour that broke the earlier run did not recur.

Zero reasoning tokens at a frontier price

Fable 5.1 reported 0 reasoning tokens across all nine tasks. Of the models in our set, that puts it in a small group — and it is the only one in that group carrying a $10-and-$50 rate card.

That shapes the economics in a way the price page hides. Reasoning tokens bill at the output rate, so a model that emits none has a bill that tracks its visible answer and nothing else. On our suite Fable 5.1 produced 1,273 output tokens across nine tasks. Gemini 3.8 Flash produced 9,168 for the same nine answers, at a seventh of the per-token price — and still landed at less than half the cost. Expensive and terse versus cheap and verbose is a real trade, and which side wins depends entirely on your prompt shape.

We have been measuring that trade all summer, in both directions: GPT-5.1-Codex-Max lists below its sibling and costs 3.09x more on volume alone, while GLM-5.3-Flash emits 1.8x the tokens of its full-size sibling and bills a tenth.

Against Astra and Opus 5

ModelScoreMeasured cost / 1kLatencyReasoning tokensList price
Claude Opus 4.89/9$4.056.1s0—
Claude Opus 59/9$5.645.3s6$5 / $25
Claude Fable 5.19/9$8.096.8s0$10 / $50
GPT-6 Astra9/9$8.195.6s47$10 / $50

Two readings. Within Anthropic's own lineup, Fable 5.1 costs twice what Opus 4.8 costs for the same 9 out of 9 and is slower than Opus 5 — on this suite, the flagship is the worst value of the three. Against OpenAI, Fable 5.1 and GPT-6 Astra land ten cents apart per thousand tasks on an identical rate card, which makes them the same purchase on this workload. We take that pair apart in the head-to-head.

Who should pay $10 and $50

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure — an empty body, a filter, a timeout — is recorded separately from a wrong answer, which is the fix Fable 5 forced on us. Cost is derived from measured token counts at the list price captured 2026-09-15, not a billing statement. Runs go through OpenRouter. Full method on the methodology page, and prices move.

What we did not measure

FAQ

How much does Claude Fable 5.1 cost? $10 per million input tokens and $50 output as of 2026-09-15. On our nine tasks that came to $8.09 per 1,000 tasks.

Is Fable 5.1 better than Fable 5? On our suite, measurably: 5.1 answered all nine tasks where 5 returned empty bodies on four of them.

How many reasoning tokens does Fable 5.1 use? Zero, across all nine tasks.

Fable 5.1 or GPT-6 Astra? On this workload they are the same purchase — both 9 out of 9, $8.09 against $8.19, on an identical $10 and $50 rate card.

Is it better value than Claude Opus 4.8? Not on these tasks. Opus 4.8 scored the same 9 out of 9 at $4.05.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.