Model Review

Claude Opus 5 Review: 9/9 at $5.64 per 1,000 Tasks (39% More Than Opus 4.8 at the Identical List Price)

We ran Claude Opus 5 through our executed coding benchmark on 2026-07-30, six days after it launched. It scored 9/9, at a measured $5.64 per 1,000 tasks, in 5.3 s average, with 6 reasoning tokens per task. Here is the part a pricing page cannot tell you: Claude Opus 4.8 carries the identical $5 / $25 list price and measured $4.05 on the same nine tasks for the same 9/9. Opus 5 costs 39% more money at an unchanged sticker price, purely because it returned more output tokens. It is faster — 5.3 s against 6.1 s. So the upgrade is a latency win and a cost increase at the same time, and nothing about the rate card says so.

Bar chart of measured cost per 1,000 tasks for six Claude models - Haiku 4.5 through Opus 5 Fast

Claude Opus 5 launched on 2026-07-24 at the same list price as Claude Opus 4.8: $5 per 1M input tokens and $25 per 1M output. Read that on a pricing page and the upgrade looks free. We ran both models on the same nine executed Python tasks, and it is not free. Opus 5 measured $5.64 per 1,000 tasks against Opus 4.8's $4.05 — 39% more, for the identical 9/9 score. The cause is not a price change. It is that Opus 5 returned more tokens to do the same work.

That is the whole reason this page exists. Everything else about Opus 5 — the 1M context window, the 128K output ceiling, the new xhigh effort level — you can read on Anthropic's own site, and we cite it below with dates. The cost-per-completed-task gap at an unchanged sticker price is something you only find by running it.

What we measured on 2026-07-30

Two Opus 5 variants went through the harness on 2026-07-30, both priced at rates captured from the live catalog the same day:

Both cleared all nine tasks with no misses. That is not a distinguishing result: 20 of the 23 models we have run scored 9/9, including every Claude model in the ladder table below. Nine short, self-contained Python functions do not separate serious 2026 coding models on correctness. What they do separate is cost and latency, and those are the two columns worth reading.

One detail worth flagging before the comparison. Anthropic ships Opus 5 with adaptive thinking on by default. We sent no effort parameter and no thinking budget — plain requests at temperature 0. Opus 5 emitted 6 reasoning tokens per task. Opus 5 Fast emitted the same 6. In practice the adaptive path looked at nine bounded function specs and declined to think about them, which is the correct call and also means this page tells you nothing about how Opus 5 behaves when it does decide to think.

Same $5 / $25, 39% more money

Claude Opus 5 and Claude Opus 4.8 list at exactly the same rate — $5 / $25 per 1M tokens. On our nine tasks they returned the same score. The measured cost per 1,000 completed tasks differs by 39%.

Model · all 9/9List price in / out per 1MMeasured cost / 1k tasksMean latencyReasoning tokens / taskPriced at
Claude Opus 4.8$5 / $25$4.056.1 s02026-07-17
Claude Opus 5$5 / $25$5.645.3 s62026-07-30
Claude Opus 5 Fast$10 / $50$10.203.4 s62026-07-30

The two pricing dates are thirteen days apart, and that is worth addressing head on, because a stale price would be a boring explanation for an interesting gap. It is not a price change. Opus 4.8 listed at $5 / $25 on 2026-07-17 and Opus 5 listed at $5 / $25 on 2026-07-30. Same rate, both dates verified against the live catalog. With the price held constant and the nine prompts held constant, a higher measured cost can only come from one place: more tokens.

We can show one side of that directly. Opus 5 consumed 896 input tokens and returned 1,850 output tokens across the nine tasks. Opus 5 Fast consumed the identical 896 input tokens — same prompts, byte for byte — and returned 1,657 output. Our run record does not store per-model token totals for the original 13-model sweep, so we cannot print Opus 4.8's split, and we are not going to back it out and present a reconstruction as a measurement. The logic stands without it: identical rate, identical input, higher measured cost, therefore more output.

What that means operationally is that "same price" and "same cost" are different claims, and only the first one appears on a rate card. A model that returns more output tokens on identical prompts is more expensive on identical prompts at an unchanged per-token rate, and output is where the money is: $25 per 1M against $5 for input. Put that inside an agent loop making twenty calls per run and the multiplier compounds the same way — the mechanism we walk through in what AI agents actually cost.

Projected out: a thousand tasks of roughly this size per day works out to about $2,060 a year on Opus 5 against about $1,480 on Opus 4.8 — a projection from our measured per-task figures at the stated rates, not a bill anyone sent us. For your own token mix, the cost calculator does the arithmetic.

Opus 5 does buy something for that money in our run: 0.8 s per call, 5.3 s against 6.1 s. Whether 0.8 s is worth 39% depends entirely on whether a person is waiting. On batch work it is worth nothing.

The whole measured Claude ladder

Six Anthropic models have now been through this harness. All six returned 9/9 with no misses. Rows are ordered by measured cost, cheapest first.

ModelScoreMeasured cost / 1k tasksMean latencyReasoning tokens / taskList price in / out per 1MPriced at
Claude Haiku 4.59/9$0.943.7 s0$1 / $52026-07-29
Claude Sonnet 59/9$1.677.2 s0$2 / $102026-07-17
Claude Sonnet 4.69/9$2.224.9 s0$3 / $152026-07-29
Claude Opus 4.89/9$4.056.1 s0$5 / $252026-07-17
Claude Opus 59/9$5.645.3 s6$5 / $252026-07-30
Claude Opus 5 Fast9/9$10.203.4 s6$10 / $502026-07-30
The measured Claude ladder: same 9/9, $0.94 to $10.20 per 1,000 tasksNine executed Python tasks, temperature 0, one scored attempt each. All six bars scored 9/9.Claude Haiku 4.5$0.94Claude Sonnet 5$1.67Claude Sonnet 4.6$2.22Claude Opus 4.8$4.05Claude Opus 5$5.64Claude Opus 5 Fast$10.20One scale throughout: 50 px per dollar. Cost is measured token counts multiplied by list price on the date in the table above, not a vendor invoice.
Chart: DataLLM Lab. Scores, latencies and token counts are measured on our executed nine-task benchmark; cost is those token counts multiplied by each model's list price on the date shown in the table above. Method: our methodology. Full run: the coding cost benchmark.

Cheapest to priciest rung is 10.9x, and the score column never moves. Opus 5 sits at 6.0x Claude Haiku 4.5's measured cost for the same nine results, and Haiku 4.5 was also 1.6 s faster. That is not an argument that Opus 5 is a worse model. It is an argument that nine bounded Python functions sit far below the level at which Opus pricing starts buying anything this harness can detect — the same conclusion the Sonnet versus Opus tier question keeps arriving at.

Note also that cost order does not predict latency order on this ladder. Sonnet 5 costs less than the older Sonnet 4.6 and is 2.3 s slower. Opus 5 costs more than Opus 4.8 and is 0.8 s faster. Cost and latency are separate axes, and inside a tier the newer model is not reliably the better one on either.

Opus 5 Fast: $10.20 for 1.9 seconds

Claude Opus 5 Fast averaged 3.4 s per task — the fifth-fastest of the 45 models we have run as of 2026-08-06, behind Llama 4 Scout at 1.5 s, GPT-5.4 mini and Gemini 3 Flash Preview at 2.3 s, and Mistral Medium 3.5 at 2.9 s. It also scored 9/9. It lists at $10 / $50 per 1M, exactly double standard Opus 5, and it measured $10.20 per 1,000 tasks.

So the trade is: 1.8x the measured cost for 1.9 seconds per call, at an identical score. State it that plainly and the decision gets easy in both directions. If a human is watching a completion render, 1.9 s off a 5.3 s wait is a real product difference. If the calls run in a queue, the extra $4.56 per 1,000 tasks buys nothing.

One asymmetry is worth naming, because it is the opposite of what the Opus 5 versus Opus 4.8 comparison showed. Opus 5 Fast returned fewer tokens than standard Opus 5 — 1,657 output against 1,850, on the identical 896 input tokens — and still cost 1.8x. Its entire premium is the doubled list rate, not verbosity. Where Opus 5's increase over Opus 4.8 came purely from token count at a flat rate, Fast's increase over Opus 5 comes purely from rate despite a lower token count. Two variants, two completely different cost mechanisms, neither visible without running both.

For context on where 3.4 s ranks in the wider field: the four fastest models on this harness are Mistral Medium 3.5 at 2.9 s, Claude Opus 5 Fast at 3.4 s, GPT-5.4 at 3.6 s and Claude Haiku 4.5 at 3.7 s. Three of those four measured under $2 per 1,000 tasks — Mistral Medium 3.5 at $0.87 priced 2026-07-17, Haiku 4.5 at $0.94 and GPT-5.4 at $1.69 both priced 2026-07-29. Speed is not what Opus 5 Fast's price is buying you relative to the field — it is buying you speed plus Opus-tier capability on the workloads this harness never reaches, which is a claim we cannot check. The AI coding ranking lays out how differently the field sorts depending on which axis you pick.

The 1M context window, and why we cannot vouch for it

The headline capability of Opus 5 is context, and we should be explicit that this is where our evidence runs out entirely.

What Anthropic states, on anthropic.com/claude/opus and in the 2026-07-24 launch announcement, read 2026-07-30:

Anthropic also positions Opus 5 as approaching Claude Fable 5's intelligence at half the price. That is their claim, not our measurement. The one part of it we can check independently is the arithmetic: Fable 5 lists at $10 / $50 per 1M and Opus 5 at $5 / $25, both captured from the live catalog on 2026-07-29 and 2026-07-30 respectively, so "half the price" is literally true on the rate card. The intelligence half we have no data on — Claude Fable 5 has never been through this harness, so we will not put it in a table next to Opus 5. What we know about it is in our Fable 5 writeup, and it is vendor-sourced.

The admission that matters. Our nine tasks are short, self-contained Python functions. The largest prompt in the set is a few hundred tokens. Total input across all nine was 896 tokens — roughly 0.09% of the 1M window. This benchmark exercises none of the capability Opus 5 is built around. We have zero first-party evidence on its long-context behaviour, and anyone telling you a nine-function coding score predicts 900,000-token retrieval quality is selling you something.

If long context is your actual use case, the honest guidance is that window size is a ceiling and not a performance guarantee. Retrieval quality degrades unevenly as a window fills, the failure mode is quiet, and it is workload-specific enough that you have to measure it on your own documents — the problem we cover in context rot. A 1M window with no beta header and no price premium removes the friction from trying; it does not tell you what you will get.

Who should move to Opus 5

Move for the context window and the latency, not for the price. If you are on Opus 4.8 and hitting the 1M window, or you need the 128K output ceiling, or 0.8 s per call matters to your product, Opus 5 is a straightforward upgrade at an unchanged per-token rate. Budget for roughly 39% more spend on comparable work anyway — that is what we measured, and the rate card will not warn you.

Do not move expecting a free upgrade. The identical $5 / $25 sticker makes it look like one. On our nine tasks it was not: same score, 39% more measured cost. If your workload is bounded generation that Opus 4.8 already handles, staying put is the cheaper answer and nothing in our data argues otherwise.

Do not run either Opus 5 variant for bounded, clearly specified code generation at all, on this evidence. Haiku 4.5 returned the same 9/9 at $0.94 and 3.7 s. That is 6.0x cheaper and 1.6 s faster than standard Opus 5 for an identical result on identical tasks. Reach for Opus when the work is genuinely ambiguous, long-context, or multi-step — the things we did not test. For the bounded tier, the cheap coding model roundup is the more relevant page.

Opus 5 Fast is a latency purchase, full stop. 1.8x the measured cost, 1.9 s saved, same score. Buy it when a person is waiting and 1.9 s per call is worth the extra $4.56 per 1,000 tasks. Otherwise do not.

The general rule. Rank by cost per completed task, not by per-token price. Two models on the same rate card can differ by 39% on the same work because verbosity is a cost input and it is not printed anywhere. Measure the models you are choosing between on your own prompts, and measure the token counts, not just the invoice total.

Run Opus 5 against Opus 4.8 on one key

One OpenAI-compatible endpoint, 300+ models, one key. Swap the model id between Opus 5, Opus 4.8 and Haiku 4.5, send your own prompts, and compare the token counts yourself.

How these numbers were produced

Nine executed Python tasks: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line.

The model receives a function signature and a prose spec. It never sees the assertions. The returned code runs against hidden assertions in an isolated python3 -I subprocess with a 12-second timeout. Every assertion passes or the task fails — no partial credit, no human grader, no LLM judge.

Temperature 0, max_tokens 4000, no effort or thinking-budget parameter set. One scored attempt per task. The harness retries only on an API error, never on a wrong answer, which is why a miss would have stayed a miss.

Cost is derived, not invoiced. It is the token counts the API reported multiplied by that model's list price on a stated date. Opus 5 and Opus 5 Fast ran on 2026-07-30 and are priced at 2026-07-30 rates. List prices move — 49 of roughly 396 catalogue models changed price in the twelve days to 2026-07-29 — so an undated cost figure is not a fact. Every number here is true as of its pricing date and should be recomputed before you act on it, a habit we argue for in LLM price volatility.

Calls go through OpenRouter's OpenAI-compatible endpoint, deliberately not through the DataLLM Lab gateway. Nothing on this page depends on our infrastructure and you do not have to be our customer to reproduce it.

On the counting. Our original sweep was 13 models run in one sitting; 10 of those 13 scored 9/9. Opus 5 and Opus 5 Fast are among ten models run later on the same harness under the same settings, which brings the total to 23, of which 20 scored 9/9. Where this page says 23 models or 20 of 23, that is the combined set. The core sweep was and remains 13.

What we did not measure

Not measured at all: long-context reasoning, multi-file refactoring, agentic or multi-turn tool use, any non-Python language, and vision. The harness is single-turn. It does not run agents and it does not use tools. Every capability Opus 5 was announced for sits in that list, including the 1M window that brought you to this page.

The effort setting is untested. We sent plain requests with no effort parameter, so the five-level control and the new xhigh mode are third-party facts here, not measurements. Opus 5's 6 reasoning tokens per task reflect what adaptive thinking chose to do on nine easy function specs, and tell you nothing about cost or quality at a higher effort level. A model that spends more reasoning tokens costs more, because reasoning bills at the output rate — so our $5.64 should be read as a floor for this model, not a typical figure.

A known scoring artefact: the 4,000-token ceiling can truncate a very verbose model mid-answer, and a truncated answer scores as a miss. That penalises verbosity, which is a real production cost but is not the same thing as being wrong. Neither Opus 5 variant was affected — both scored 9/9 — but the ceiling is part of the harness.

Nine tasks is nine data points. A 9/9 is a clean result on this set, not a claim about a distribution. We did not run it twice, we did not vary the prompts, and we did not buy any model a retry.

Not tested, and never claimed as ours: Claude Fable 5, GPT-5 nano, Ollama or any locally-run model, Grok 3, Grok 4, Grok 4.20, grok-build-0.1, VibeThinker, and any Gemini other than 3.6 Flash and 3.1 Pro. If a first-party number for any of those appears anywhere on this site, it is an error.

FAQ

What is the Claude Opus 5 context window and output limit?

Anthropic states a 1,000,000-token context window with no beta header required and no long-context price premium, and up to 128,000 output tokens per synchronous request (anthropic.com/claude/opus and the 2026-07-24 launch announcement, read 2026-07-30). The 1M figure covers input plus output combined; the 128K figure is the separate cap on output alone. We have not verified either number by measurement — our nine tasks used 896 input tokens in total, about 0.09% of the window. See the Claude context window reference for the per-model table.

Is Claude Opus 5 more expensive than Opus 4.8?

Not on list price — both are $5 / $25 per 1M tokens, verified 2026-07-30 and 2026-07-17 respectively. But on our measured cost per completed task, yes: $5.64 per 1,000 tasks for Opus 5 against $4.05 for Opus 4.8, both scoring 9/9 on the same nine tasks. That is 39% more at an identical rate, and the cause is token count — Opus 5 returned more output on the same prompts. Opus 5 was faster, 5.3 s against 6.1 s.

Is Claude Opus 5 Fast worth double the price?

It bought 1.9 seconds in our run. Opus 5 Fast averaged 3.4 s at a measured $10.20 per 1,000 tasks; standard Opus 5 averaged 5.3 s at $5.64. Both were priced at 2026-07-30 rates. Both scored 9/9. That is 1.8x the cost for 1.9 s per call at an identical score, so it is worth it only where a human is waiting on the response. Its list price is $10 / $50 against $5 / $25, and interestingly it returned fewer tokens than standard Opus 5 — the entire premium is the rate, not verbosity.

Did Opus 5 beat the cheaper Claude models on your benchmark?

No. All six Claude models we have run scored 9/9: Haiku 4.5, Sonnet 5, Sonnet 4.6, Opus 4.8, Opus 5 and Opus 5 Fast. Haiku 4.5 did it at $0.94 per 1,000 tasks and 3.7 s (priced 2026-07-29), against Opus 5's $5.64 and 5.3 s (priced 2026-07-30). Nine short Python functions are inside the competent range of every serious 2026 coding model — 20 of the 23 models we have run scored 9/9 — so a tie means the test did not reach the tiers, not that the tiers are identical.

Is Claude Opus 5 as good as Claude Fable 5?

We have no idea, and we will not pretend otherwise: Claude Fable 5 has never been through our harness. Anthropic positions Opus 5 as approaching Fable 5's intelligence at half the price. The pricing half of that claim checks out — Fable 5 lists at $10 / $50 per 1M and Opus 5 at $5 / $25, captured 2026-07-29 and 2026-07-30. The intelligence half is Anthropic's claim, unverified by us.

Is this measured cost the same as my bill?

No. We take the token counts the API reported and multiply by each model's list price on a stated date — 2026-07-30 for both Opus 5 variants. It is a derived measured cost, not a vendor invoice. Your bill will differ with prompt length, caching, discounts, provider routing, effort level and any price change since that date. The ratios between models are the durable part; the absolute dollars are not. Full rate card in the Claude API pricing guide.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.