Model Reviews

Claude Opus 5.5 Review: 9/9 at $4.03 (Two Cents Under Opus 4.8, Twice Sonnet 5.5)

Claude Opus 5.5 scored 9 out of 9 on our executed Python benchmark at $4.03 per 1,000 tasks, priced 2026-10-02, with a 4.8-second mean and 33 reasoning tokens per call. The surprising number is the one beside it: Claude Opus 4.8 measured $4.05 in July, priced 2026-07-17. Anthropic's new, lower Opus price does not take the measured bill somewhere new — it takes it back, almost to the cent, to where Opus 4.8 already was before Opus 5 pushed it up. And Claude Sonnet 5.5 cleared the same nine tasks for $1.95, also priced 2026-10-02, reading exactly the same 914 input tokens.

DataLLM Lab article cover: Claude Opus 5.5 Review: 9/9 at $4.03 (Two Cents Under Opus 4.8, Twice Sonnet 5.5)

A price cut on a flagship reads like a discount. Measured on the same nine tasks, this one reads more like a correction.

The result

MetricClaude Opus 5.5
Score9/9
Measured cost / 1,000 tasks$4.03 (priced 2026-10-02)
Mean latency4.8s
Reasoning tokens per call33
Tokens across the suite914 in / 1,630 out
List price in / out$4 / $20 per 1M (2026-10-02)
Context window1,000,000
Measured2026-10-02

Of the 75 models in our set that had scored 9 out of 9 as of 2026-10-02, Opus 5.5 ranks 48th cheapest and 20th fastest. It is also terse: 33 reasoning tokens per call is close to not thinking at all on problems this size, which is the right call for them.

For context on the vendor's side, and labelled as third-party: TechCrunch reported on 2026-09-22 that Anthropic released Opus 5.5 that day, at a lower price than Opus 5, and that Anthropic positions it at Claude Fable 5.1 level on most tasks (read 2026-10-02). Our suite can test the price half of that pitch. It cannot test the Fable half, for reasons we get to below.

Against Opus 5 and Opus 4.8

ModelScoreMeasured cost / 1kPriced onList in / outLatencyReasoning tokensTokens in / out
Claude Opus 5.59/9$4.032026-10-02$4 / $204.8s33914 / 1,630
Claude Opus 59/9$5.642026-07-30$5 / $255.3s6896 / 1,850
Claude Opus 4.89/9$4.052026-07-17see note6.1s0not stored
Claude Sonnet 5.59/9$1.952026-10-02$2 / $103.4s31914 / 1,569

Against Claude Opus 5 the story is clean, because both runs stored their token counts. Opus 5.5 costs $1.61 less per thousand tasks ($5.64 − $4.03), which is 28.5% less (1.61 / 5.64). Two things produce that. The rate card dropped $1 per million on input and $5 on output, so every token is 0.8 of its old price (4 / 5, and 20 / 25). And Opus 5.5 wrote 220 fewer output tokens for the same nine answers (1,850 − 1,630). Most of the saving is the price; a real slice of it is the model being less wordy.

Against Claude Opus 4.8 the picture is less flattering. Opus 4.8 measured $4.05 in our July core-13 sweep, priced 2026-07-17. Opus 5.5 lands two cents below that ($4.05 − $4.03). When we reviewed Opus 5 we found it cost more than Opus 4.8 at an identical list price because it emitted more tokens. Opus 5.5 is, measured on this suite, Anthropic buying that back. One honest caveat: the Opus 4.8 entry predates the field where we store run-date prices and token counts, so we cannot split its $4.05 into price and tokens the way we can for Opus 5. Its list price today is $5 / $25 per million. The tier-by-tier history is in our Sonnet versus Opus comparison.

Latency moved the right way across the line: 6.1 seconds for Opus 4.8, 5.3 for Opus 5, 4.8 for Opus 5.5. Reasoning tokens moved the other way, from 0 to 6 to 33 per call — still small enough that it barely registers in the bill.

Sonnet 5.5: same score, half the bill

This is the comparison that decides most purchases. Claude Sonnet 5.5 went through the same harness on the same day:

So the two models did essentially the same amount of work, and the cost gap is almost entirely the rate card: $4.03 against $1.95, both priced 2026-10-02, or 2.07 times (4.03 / 1.95). Sonnet 5.5 was also faster, at 3.4 seconds against 4.8, and ranks 9th fastest of the 75 models at 9/9. On this suite there is nothing Opus 5.5 did that Sonnet 5.5 did not, and Sonnet did it quicker. We say this with the usual limit attached: a ceiling-hit suite shows a tie, not equivalence.

Opus 5.5 lands two cents under Opus 4.8. Sonnet 5.5 ties it for less.Measured cost per 1,000 tasks, all five 9/9 on the same nine executed Python tasks. Dates are price-capture dates.Sonnet 5.5 · 10-02$1.95Opus 5.5 · 10-02$4.03Opus 4.8 · 07-17$4.05Opus 5 · 07-30$5.64Fable 5.1 · 09-15$8.09JEV ROUTER PROBE, 2026-10-02 · 9 TASKS · API-REPORTED COSTWhole billparse_csv_line on Opus 5.5 · $0.0051288 other tasks · $0.003969$0.009097One call of nine was 56% of the router's bill.Top panel scale: 66 px per dollar ($4.03 × 66 = 265.98 px). Every top bar is its dollar value × 66.Bottom panel scale: 60,000 px per dollar ($0.005128 × 60,000 = 307.68 px; $0.003969 × 60,000 = 238.14 px).
Top: derived cost at list price. Bottom: cost the API reported per call. Different methods, so the panels use different scales and should not be compared bar to bar.

One router call, 56% of the bill

On 2026-10-02 we also ran our nine tasks through typesafe/jev-router. Labelled as third-party, from OpenRouter's documentation and launch material (read 2026-10-02): the router picks a model and a reasoning effort per request. Here is what it actually picked:

TaskServed byAPI-reported costLatencyPass
two_sumgpt-6-luna$0.0000463.3syes
valid_parenthesesdeepseek-v4.1-flash$0.00017341.8syes
merge_intervalsgpt-6.1-sol$0.0009284.4syes
roman_to_intdeepseek-v4.1-flash$0.00019592.3syes
lcs_lengpt-6-luna$0.00006895.9syes
flattengpt-6-luna$0.00009344.7syes
top_k_wordsdeepseek-v4.1-flash$0.00037322.0syes
token_bucketgpt-6.1-sol$0.002097.7syes
parse_csv_lineclaude-opus-5.5$0.0051284.9syes

The router passed all nine for a summed $0.009097, which is $1.01 per 1,000 tasks. Opus 5.5 was used exactly once, and that one call was 56% of the total. The other eight tasks together came to $0.003969 ($0.009097 − $0.005128).

The router's instinct was not wrong about which task is hard. In this sweep, Codestral 2508, Qwen3 Coder Plus and Qwen3 Coder Flash each dropped parse_csv_line and nothing else; GLM-5.3 Prime dropped it alongside token_bucket; and Cohere Command A Plus was excluded from scoring, with no valid score, because its parse_csv_line response came back empty at the token limit. Quoted CSV with escapes is where our suite bites.

But on our evidence the escalation was not necessary. GPT-6 Luna, which the router used for three of the easy tasks, cleared all nine including parse_csv_line on its own, at $0.16 per 1,000 tasks priced 2026-10-02. The router paid flagship rates for insurance it did not need here. Whether that insurance pays on your workload is the right question to ask of any router; our probe is one pass, and its cost column is what the API reported per call rather than our derived figure, so do not set $1.01 directly against the derived numbers above.

Is it worth $4.03?

On nine self-contained Python functions, no. As of 2026-10-02 the cheapest and fastest model to score 9 out of 9 is Upstage Solar Mini 4, at $0.03 (priced 2026-10-02) and 2 seconds — about 134 times cheaper (4.03 / 0.03). Our suite cannot distinguish a frontier model from a competent small one, and pretending otherwise would misuse it. The cheap coding roundup covers that end of the field.

What it can say about Anthropic's own line-up is narrower and firmer. Of the three Opus models in the table above, it is the cheapest, by two cents over Opus 4.8. It clears the suite for about half of what Claude Fable 5.1 measured ($8.09, priced 2026-09-15; 4.03 / 8.09 is 0.498). Whether it matches Fable on hard work, as Anthropic claims, both models hit our ceiling, so we have no data either way. And Sonnet 5.5 matched it here for $1.95. If you are buying Opus 5.5, buy it for work you have measured Sonnet 5.5 failing on.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why a model with an empty response is excluded rather than marked wrong. Cost is derived — measured input and output token counts multiplied by the list price captured on the run date — not a billing statement, and prices move. The router probe is the exception: its cost column is the per-call cost the API reported. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

What we did not measure

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.