Model Reviews

Claude Sonnet 5.5 Review: 3.4 Seconds, 31 Reasoning Tokens (9/9)

Claude Sonnet 5.5 answered our nine coding tasks in a 3.4-second mean — less than half the 7.2 seconds its predecessor Sonnet 5 took on the same suite. It scored 9 out of 9, emitted only 31 reasoning tokens per call, and ranked 9th fastest of the 75 models that had cleared the suite as of 2026-10-02. On cost it is unremarkable: $1.95 per 1,000 tasks at the $2 in / $10 out list price captured 2026-10-02, which puts it 34th cheapest. Against GPT-6 Sol on the identical $2/$10 rate card, Sonnet 5.5 was 8 cents cheaper and 1.8 seconds faster. Against OpenAI's newer GPT-6.1 Sol, it lost on cost.

DataLLM Lab article cover: Claude Sonnet 5.5 Review: 3.4 Seconds, 31 Reasoning Tokens (9/9)

Anthropic's launch pitch for Sonnet 5.5, as reported by Help Net Security and The New Stack (third-party coverage, read 2026-10-02), is speed without a price increase. That is a claim we can test directly, so we did.

The result

MetricClaude Sonnet 5.5
Score9/9
Measured cost / 1,000 tasks$1.95 (priced 2026-10-02)
Mean latency3.4s
Reasoning tokens per call31
Tokens across the suite914 in / 1,569 out
List price in / out$2 / $10 per 1M (2026-10-02)
Context window1,000,000
Measured2026-10-02

Of the 75 models that had scored 9 out of 9 as of 2026-10-02, Sonnet 5.5 ranked 9th on latency and 34th on cost. That is the profile of a mid-priced model that does not waste time: 31 reasoning tokens per call means it barely deliberates before writing code, and on nine self-contained functions it does not need to.

For scale, the fastest and cheapest model to clear the suite as of 2026-10-02 is Upstage's Solar Mini 4, at 2 seconds and $0.03 per 1,000 tasks priced 2026-10-02. Sonnet 5.5 costs 65 times as much (1.95 ÷ 0.03) for the same score. Nobody buys Sonnet for this suite; the question is what the generation change bought.

Against Sonnet 5: faster, and we cannot say if dearer

Claude Sonnet 5Claude Sonnet 5.5
Score9/99/9
Mean latency7.2s3.4s
Reasoning tokens per call031
Measured cost / 1,000 tasks$1.67 (priced 2026-07-17)$1.95 (priced 2026-10-02)
Tokens across the suitenot stored914 in / 1,569 out
Rank among the 75 at 9/932nd cheapest · 39th fastest34th cheapest · 9th fastest

The speed claim holds on our suite. Mean latency fell from 7.2 seconds to 3.4, a 3.8-second drop, and Sonnet 5.5 moved from 39th fastest to 9th. It got there while spending 31 reasoning tokens per call where Sonnet 5 spent none, so the gain is not from thinking less. Whatever changed is in serving or generation speed, which we observe but cannot attribute.

The cost comparison is weaker than it looks, and we want to be precise about why. Sonnet 5 came from our original core-13 sweep, before we stored per-run token counts or the run-date price. Its $1.67 was derived at the list price on 2026-07-17, and our data does not record what that price was. Today Sonnet 5 lists at $2 / $10 per 1M, the same as Sonnet 5.5, and launch coverage says Anthropic kept the price unchanged. If that held in July, Sonnet 5.5's 28-cent increase ($1.95 − $1.67) would be pure token volume. We cannot verify that it did, and without Sonnet 5's token counts we cannot decompose the gap. So we report it as two numbers priced on two dates, not as a regression.

Within the family, Sonnet 5.5 also beat Claude Opus 5.5 on both axes on this suite: Opus 5.5 scored the same 9/9 at $4.03 per 1,000 tasks (priced 2026-10-02) and 4.8 seconds, on a $4 / $20 list that is exactly twice Sonnet's. Their output volumes were nearly identical, 1,569 tokens against 1,630, so the cost gap is almost entirely the rate card. Our Sonnet vs Opus tier comparison covers the earlier generations.

Against GPT-6 Sol on the same rate card

Four models in our October data were measured on the same day at the same $2 in / $10 out list: Sonnet 5.5 and three OpenAI Sol variants. When the rate card is identical, every cent of difference is token count.

ModelScoreCost / 1k tasks (priced 2026-10-02)LatencyReasoning / callTokens in / out
Claude Sonnet 5.59/9$1.953.4s31914 / 1,569
GPT-6 Sol9/9$2.035.2s85604 / 1,704
GPT-6.1 Sol9/9$1.615.9s48604 / 1,329
GPT-6 Sol Pro9/9$7.455.9s15816,898 / 3,322
Same score, same rate card: Sonnet 5.5 is the quick one, not the cheap oneAll five scored 9/9 on the identical nine executed Python tasks.MEASURED COST PER 1,000 TASKS · $2 IN / $10 OUT · PRICED 2026-10-02GPT-6.1 Sol$1.61Claude Sonnet 5.5$1.95GPT-6 Sol$2.03GPT-6 Sol Pro$7.45MEAN LATENCY, SECONDSClaude Sonnet 5.53.4sGPT-6 Sol5.2sGPT-6.1 Sol5.9sGPT-6 Sol Pro5.9sClaude Sonnet 57.2sScales: cost 70 px per dollar (bar = cost × 70); latency 60 px per second (bar = seconds × 60).Sonnet 5 appears in latency only: its cost was priced on 2026-07-17, not 2026-10-02.
Fastest of these five on latency; second cheapest of the four priced on 2026-10-02.

Against GPT-6 Sol, Sonnet 5.5 wins narrowly on cost and clearly on speed. It was 8 cents per 1,000 tasks cheaper ($2.03 − $1.95, both priced 2026-10-02) and 1.8 seconds faster. The mechanics are worth seeing. The same nine prompts counted as 914 input tokens on Sonnet 5.5 and 604 on GPT-6 Sol — 310 more for Claude before either model writes a word. The same 914 appears for Opus 5.5 and Claude Fable 5.1, and 604 for every GPT-6 variant without a Pro suffix, which points to tokenizer differences rather than anything either model chose to do; that is our inference, not something we measured directly. Sonnet 5.5 claws it back on output: 1,569 tokens against Sol's 1,704, 135 fewer. Output costs five times input on this rate card ($10 ÷ $2), so writing less beats reading more.

Against GPT-6.1 Sol, it loses on cost. GPT-6.1 Sol cost $1.61 (priced 2026-10-02), 34 cents less than Sonnet 5.5, because it read the cheaper 604-token version of the prompts and wrote 240 fewer output tokens (1,569 − 1,329). It was also 2.5 seconds slower. If you are choosing on this rate card purely on measured cost, GPT-6.1 Sol is the better buy on our data; if you are choosing on latency, Sonnet 5.5 is.

GPT-6 Sol Pro is the outlier, at $7.45 per 1,000 tasks (priced 2026-10-02) for the same 9/9. It consumed 16,898 input tokens on the same nine prompts that Sol read as 604 — the same hidden-prompt pattern we documented when GPT-6 Astra and Astra Pro landed. A price page cannot show you that; a token count can.

Is it worth $1.95?

On this suite, the honest answer is that it is worth what any 9/9 model is worth, and 74 other models also scored 9/9 as of 2026-10-02. Nine self-contained Python functions cannot separate a frontier model from a competent small one. Our cheap coding roundup lists models that clear the same bar for a few cents.

What the suite can tell you is narrower and still useful. If you already run Sonnet and latency matters — interactive coding, an agent loop where every turn waits on the model — the 5.5 upgrade more than halved mean latency for us and kept the score. If cost is the constraint, it does not improve the bill on our data, and GPT-6.1 Sol undercuts it on the identical list price. For how Claude bills more generally, see our Claude API pricing guide.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, so a model that times out at the provider is excluded rather than scored as wrong. Cost is derived — measured input and output token counts multiplied by the list price captured on the stated date — not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

What we did not measure

FAQ

How much does Claude Sonnet 5.5 cost? $2 per million input tokens and $10 output as of 2026-10-02. On our nine tasks that worked out to $1.95 per 1,000 tasks at that price.

Is Claude Sonnet 5.5 faster than Sonnet 5? On our suite, yes: a 3.4-second mean against 7.2 seconds, with the same 9/9 score.

Claude Sonnet 5.5 or GPT-6 Sol? Same list price, same score. Sonnet 5.5 was 8 cents per 1,000 tasks cheaper (priced 2026-10-02) and 1.8 seconds faster.

Is it the cheapest model on the $2/$10 rate card? Not among the four we measured there on 2026-10-02: GPT-6.1 Sol was cheaper at $1.61.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.