Claude Opus 5.5 Review: 9/9 at $4.03 (Two Cents Under Opus 4.8, Twice Sonnet 5.5)
Claude Opus 5.5 scored 9 out of 9 on our executed Python benchmark at $4.03 per 1,000 tasks, priced 2026-10-02, with a 4.8-second mean and 33 reasoning tokens per call. The surprising number is the one beside it: Claude Opus 4.8 measured $4.05 in July, priced 2026-07-17. Anthropic's new, lower Opus price does not take the measured bill somewhere new — it takes it back, almost to the cent, to where Opus 4.8 already was before Opus 5 pushed it up. And Claude Sonnet 5.5 cleared the same nine tasks for $1.95, also priced 2026-10-02, reading exactly the same 914 input tokens.
A price cut on a flagship reads like a discount. Measured on the same nine tasks, this one reads more like a correction.
The result
| Metric | Claude Opus 5.5 |
|---|---|
| Score | 9/9 |
| Measured cost / 1,000 tasks | $4.03 (priced 2026-10-02) |
| Mean latency | 4.8s |
| Reasoning tokens per call | 33 |
| Tokens across the suite | 914 in / 1,630 out |
| List price in / out | $4 / $20 per 1M (2026-10-02) |
| Context window | 1,000,000 |
| Measured | 2026-10-02 |
Of the 75 models in our set that had scored 9 out of 9 as of 2026-10-02, Opus 5.5 ranks 48th cheapest and 20th fastest. It is also terse: 33 reasoning tokens per call is close to not thinking at all on problems this size, which is the right call for them.
For context on the vendor's side, and labelled as third-party: TechCrunch reported on 2026-09-22 that Anthropic released Opus 5.5 that day, at a lower price than Opus 5, and that Anthropic positions it at Claude Fable 5.1 level on most tasks (read 2026-10-02). Our suite can test the price half of that pitch. It cannot test the Fable half, for reasons we get to below.
Against Opus 5 and Opus 4.8
| Model | Score | Measured cost / 1k | Priced on | List in / out | Latency | Reasoning tokens | Tokens in / out |
|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 9/9 | $4.03 | 2026-10-02 | $4 / $20 | 4.8s | 33 | 914 / 1,630 |
| Claude Opus 5 | 9/9 | $5.64 | 2026-07-30 | $5 / $25 | 5.3s | 6 | 896 / 1,850 |
| Claude Opus 4.8 | 9/9 | $4.05 | 2026-07-17 | see note | 6.1s | 0 | not stored |
| Claude Sonnet 5.5 | 9/9 | $1.95 | 2026-10-02 | $2 / $10 | 3.4s | 31 | 914 / 1,569 |
Against Claude Opus 5 the story is clean, because both runs stored their token counts. Opus 5.5 costs $1.61 less per thousand tasks ($5.64 − $4.03), which is 28.5% less (1.61 / 5.64). Two things produce that. The rate card dropped $1 per million on input and $5 on output, so every token is 0.8 of its old price (4 / 5, and 20 / 25). And Opus 5.5 wrote 220 fewer output tokens for the same nine answers (1,850 − 1,630). Most of the saving is the price; a real slice of it is the model being less wordy.
Against Claude Opus 4.8 the picture is less flattering. Opus 4.8 measured $4.05 in our July core-13 sweep, priced 2026-07-17. Opus 5.5 lands two cents below that ($4.05 − $4.03). When we reviewed Opus 5 we found it cost more than Opus 4.8 at an identical list price because it emitted more tokens. Opus 5.5 is, measured on this suite, Anthropic buying that back. One honest caveat: the Opus 4.8 entry predates the field where we store run-date prices and token counts, so we cannot split its $4.05 into price and tokens the way we can for Opus 5. Its list price today is $5 / $25 per million. The tier-by-tier history is in our Sonnet versus Opus comparison.
Latency moved the right way across the line: 6.1 seconds for Opus 4.8, 5.3 for Opus 5, 4.8 for Opus 5.5. Reasoning tokens moved the other way, from 0 to 6 to 33 per call — still small enough that it barely registers in the bill.
Sonnet 5.5: same score, half the bill
This is the comparison that decides most purchases. Claude Sonnet 5.5 went through the same harness on the same day:
- Input tokens: 914 for both. Same prompts, same tokenizer, same count to the token.
- Output tokens: 1,630 for Opus, 1,569 for Sonnet. 61 apart across nine tasks (1,630 − 1,569).
- Reasoning: 33 against 31 per call.
- List price: $4 / $20 against $2 / $10, both captured 2026-10-02. Exactly double.
So the two models did essentially the same amount of work, and the cost gap is almost entirely the rate card: $4.03 against $1.95, both priced 2026-10-02, or 2.07 times (4.03 / 1.95). Sonnet 5.5 was also faster, at 3.4 seconds against 4.8, and ranks 9th fastest of the 75 models at 9/9. On this suite there is nothing Opus 5.5 did that Sonnet 5.5 did not, and Sonnet did it quicker. We say this with the usual limit attached: a ceiling-hit suite shows a tie, not equivalence.
One router call, 56% of the bill
On 2026-10-02 we also ran our nine tasks through typesafe/jev-router. Labelled as third-party, from OpenRouter's documentation and launch material (read 2026-10-02): the router picks a model and a reasoning effort per request. Here is what it actually picked:
| Task | Served by | API-reported cost | Latency | Pass |
|---|---|---|---|---|
| two_sum | gpt-6-luna | $0.000046 | 3.3s | yes |
| valid_parentheses | deepseek-v4.1-flash | $0.0001734 | 1.8s | yes |
| merge_intervals | gpt-6.1-sol | $0.000928 | 4.4s | yes |
| roman_to_int | deepseek-v4.1-flash | $0.0001959 | 2.3s | yes |
| lcs_len | gpt-6-luna | $0.0000689 | 5.9s | yes |
| flatten | gpt-6-luna | $0.0000934 | 4.7s | yes |
| top_k_words | deepseek-v4.1-flash | $0.0003732 | 2.0s | yes |
| token_bucket | gpt-6.1-sol | $0.00209 | 7.7s | yes |
| parse_csv_line | claude-opus-5.5 | $0.005128 | 4.9s | yes |
The router passed all nine for a summed $0.009097, which is $1.01 per 1,000 tasks. Opus 5.5 was used exactly once, and that one call was 56% of the total. The other eight tasks together came to $0.003969 ($0.009097 − $0.005128).
The router's instinct was not wrong about which task is hard. In this sweep, Codestral 2508, Qwen3 Coder Plus and Qwen3 Coder Flash each dropped parse_csv_line and nothing else; GLM-5.3 Prime dropped it alongside token_bucket; and Cohere Command A Plus was excluded from scoring, with no valid score, because its parse_csv_line response came back empty at the token limit. Quoted CSV with escapes is where our suite bites.
But on our evidence the escalation was not necessary. GPT-6 Luna, which the router used for three of the easy tasks, cleared all nine including parse_csv_line on its own, at $0.16 per 1,000 tasks priced 2026-10-02. The router paid flagship rates for insurance it did not need here. Whether that insurance pays on your workload is the right question to ask of any router; our probe is one pass, and its cost column is what the API reported per call rather than our derived figure, so do not set $1.01 directly against the derived numbers above.
Is it worth $4.03?
On nine self-contained Python functions, no. As of 2026-10-02 the cheapest and fastest model to score 9 out of 9 is Upstage Solar Mini 4, at $0.03 (priced 2026-10-02) and 2 seconds — about 134 times cheaper (4.03 / 0.03). Our suite cannot distinguish a frontier model from a competent small one, and pretending otherwise would misuse it. The cheap coding roundup covers that end of the field.
What it can say about Anthropic's own line-up is narrower and firmer. Of the three Opus models in the table above, it is the cheapest, by two cents over Opus 4.8. It clears the suite for about half of what Claude Fable 5.1 measured ($8.09, priced 2026-09-15; 4.03 / 8.09 is 0.498). Whether it matches Fable on hard work, as Anthropic claims, both models hit our ceiling, so we have no data either way. And Sonnet 5.5 matched it here for $1.95. If you are buying Opus 5.5, buy it for work you have measured Sonnet 5.5 failing on.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why a model with an empty response is excluded rather than marked wrong. Cost is derived — measured input and output token counts multiplied by the list price captured on the run date — not a billing statement, and prices move. The router probe is the exception: its cost column is the per-call cost the API reported. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
What we did not measure
- Hard work. Opus 5.5, Opus 5, Opus 4.8, Sonnet 5.5 and Fable 5.1 all scored 9 out of 9. A suite every one of them maxes out cannot rank them on capability, and it cannot test the Fable-level claim at all.
- Agentic and multi-turn sessions, where Opus is usually bought. Nine single-turn functions say nothing about a model running a long task with tools.
- The 1,000,000-token context. Our prompts are short.
- Prompt caching. No prompt repeats in our suite, so a cached workload could cost very differently.
- Reasoning-effort settings. We did not vary them. The router can; we cannot say what effort it used on the Opus call.
- Repeat runs. One scored attempt per task. A two-cent gap to Opus 4.8, across different sweeps and run dates, is within what a re-run could move.
- Router consistency. One pass. The router may send parse_csv_line somewhere else tomorrow.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab