Claude Haiku 4.5 Review: 9/9 at $0.94 per 1,000 Tasks (4.3x Cheaper Than Opus 4.8, Same Score)
We ran Claude Haiku 4.5 through our executed coding benchmark on 2026-07-29. It scored 9/9, at a measured $0.94 per 1,000 tasks, in 3.7 s average, with zero reasoning tokens. On the same nine tasks, Claude Opus 4.8 also scored 9/9 — at $4.05, or 4.3x more. Claude Sonnet 5 also scored 9/9 — at $1.67, or 1.8x more. The entire Claude ladder we have run went 9/9 here, so moving up it bought nothing this harness could measure and cost up to 4.3x. That result is real and it is narrow: nine short, self-contained Python functions are exactly the workload a small model handles well, and exactly the workload the bigger tiers were not built for.
Most model reviews restate a vendor announcement. This one reports a run. Claude Haiku 4.5 went through the same executed nine-task Python harness we have now put 21 models through, on 2026-07-29, and the result is unusual enough to state before anything else: the cheap tier scored exactly what the expensive tier scored.
Read this first: what nine short functions can prove
This harness runs nine short, self-contained Python functions. That is the whole scope. The model gets a signature and a prose spec, returns code, and the code is executed against assertions it never sees.
Nine bounded functions are precisely the workload a small model handles well. It is also precisely not the workload the larger tiers exist for. Anthropic sells Sonnet and Opus for long context, sustained multi-step reasoning, agentic tool use and ambiguous specifications. This harness touches none of those. It is single-turn, Python-only, and every task fits comfortably inside a few hundred tokens.
So the honest claim on this page is narrow, and we think it is worth making anyway: for bounded, clearly specified code generation, Haiku 4.5 is the value pick in the Claude family by a wide margin. If your workload looks like that, this is directly relevant. If it does not, treat the result as a reason to test the cheap tier on your own tasks rather than as permission to switch. The full boundary list is in what we did not measure, and the harness design is on the methodology page.
We put this above the numbers because a cheap-tier recommendation is exactly where a benchmark like ours goes wrong if you read it too widely.
Four Claude tiers, one score
Four Anthropic models have now been through this harness. All four returned 9/9 with no misses. Rows are ordered by measured cost, cheapest first.
| Model | Score | Measured cost / 1k tasks | Mean latency | Reasoning tokens | List price in / out per 1M | Priced at |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | 9/9 | $0.94 | 3.7 s | 0 | $1 / $5 | 2026-07-29 |
| Claude Sonnet 5 | 9/9 | $1.67 | 7.2 s | 0 | $2 / $10 | 2026-07-17 |
| Claude Sonnet 4.6 | 9/9 | $2.22 | 4.9 s | 0 | $3 / $15 | 2026-07-29 |
| Claude Opus 4.8 | 9/9 | $4.05 | 6.1 s | 0 | $5 / $25 | 2026-07-17 |
Two of these rows were priced on 2026-07-17 and two on 2026-07-29, because Sonnet 5 and Opus 4.8 ran in our original 13-model sweep and Haiku 4.5 and Sonnet 4.6 ran later on the same harness. That mixture would matter if Anthropic had moved prices in between. It did not: 49 of roughly 396 catalogue models changed price in those twelve days, and none of these four were among them. The list rates in the table are the same on both dates.
The gap between the cheapest and priciest rung is 4.3x, and the score column does not move. That is the finding. It is not that Opus 4.8 is a weak model — it is that nine short Python functions sit far below the level at which its price starts buying anything we can detect.
The arithmetic scales the way you would expect. A thousand tasks of roughly this size per day, at these rates, projects to about $340 a year on Haiku 4.5 against about $1,480 on Opus 4.8 — a projection from measured per-task numbers, not a bill anyone sent us. Put either inside an agent loop that makes twenty calls per run and the same multiple applies again, which is the mechanism described in AI agent traps.
Zero reasoning tokens, and why it shows up twice
Haiku 4.5 emitted 0 reasoning tokens across all nine tasks. It answers directly. That single property explains both halves of its result — it is cheap because reasoning tokens bill at the output rate, and it is fast because it does not spend wall-clock time producing them.
The contrast case on the same harness is Gemini 3.1 Pro. It also scored 9/9. It averaged 1,094 reasoning tokens per call, the highest of any model we have run, and its measured cost was $14.70 per 1,000 tasks priced 2026-07-29. That is a 15.6x gap for an identical score. Same nine problems, same grader, same settings.
There is a second-order point here that a per-token price list cannot show you. A model that emits reasoning tokens has variable cost per call, and the variance tracks how awkward your prompt is rather than how long it is. On our harness, Gemini 3.6 Flash spent 341 reasoning tokens on two_sum and 2,615 on parse_csv_line — 7.7x, on nine tasks that all look similar from the outside. A model at zero has no such spread. You can budget it.
Seven of the 21 models we have run emitted zero reasoning tokens, and all seven scored 9/9. Sorted by measured cost:
| Model · all 9/9, all 0 reasoning tokens | Measured cost / 1k tasks | Mean latency | Priced at |
|---|---|---|---|
| Qwen3 Coder Next | $0.10 | 7.0 s | 2026-07-17 |
| Mistral Medium 3.5 | $0.87 | 2.9 s | 2026-07-17 |
| Claude Haiku 4.5 | $0.94 | 3.7 s | 2026-07-29 |
| Claude Sonnet 5 | $1.67 | 7.2 s | 2026-07-17 |
| GPT-5.4 | $1.69 | 3.6 s | 2026-07-29 |
| Claude Sonnet 4.6 | $2.22 | 4.9 s | 2026-07-29 |
| Claude Opus 4.8 | $4.05 | 6.1 s | 2026-07-17 |
Four of those seven are Claude models. Whatever Anthropic is doing with the direct-answer path, it applies across the ladder on this workload, which means none of the four carries a thinking-token surcharge on top of its list rate — the surcharge that put Gemini 3.1 Pro at $14.70 for the same 9/9.
Third-fastest of the 21 models we have run
Haiku 4.5 averaged 3.7 s per task. Across all 21 models on this harness, only two were faster: Mistral Medium 3.5 at 2.9 s and GPT-5.4 at 3.6 s. Third place — and it is not the cheapest of the three: Mistral Medium 3.5 beats it on both axes, 2.9 s at $0.87 per 1,000 tasks priced 2026-07-17 against Haiku 4.5's $0.94 priced 2026-07-29. Against GPT-5.4, the other model in that top three, Haiku 4.5 gives up 0.1 s and saves $0.75 per 1,000 tasks, both priced 2026-07-29.
Within the Claude family it is the fastest by a clear margin: 3.7 s against Sonnet 4.6 at 4.9 s, Opus 4.8 at 6.1 s and Sonnet 5 at 7.2 s. Note that the speed order inside the family is not the price order — Sonnet 4.6 costs more than Sonnet 5 and is faster than it. Latency and cost are separate axes and they do not agree, which is the same disagreement we found across the whole field in the AI coding ranking.
Where 3.7 s matters is anywhere a human is watching: inline completions, a chat that has to feel responsive, a lint-and-fix loop. Where it does not matter is batch work, and if your workload is batch you should be reading the cost column only. For a broader view of the sub-$1 tier, see the cheap coding model roundup.
Sonnet 5 against Sonnet 4.6: the newer tier is the cheaper one
One result in the ladder table is worth pulling out because it runs against the usual assumption that a newer model costs more. Claude Sonnet 5 measured $1.67 per 1,000 tasks; the older Claude Sonnet 4.6 measured $2.22. Both scored 9/9. Both emitted zero reasoning tokens.
The reason is the list price, not the token counts: Sonnet 5 lists at $2 / $10 per 1M and Sonnet 4.6 at $3 / $15, both captured 2026-07-29. On this workload the newer mid tier is a straight upgrade on cost per completed task — you pay less and score the same.
The one thing Sonnet 4.6 wins here is latency, 4.9 s against 7.2 s. If you are on Sonnet 4.6 today and your workload looks like ours, the move to Sonnet 5 saves money and costs you about 2.3 s per call. That is a real trade, not a free win, and which side of it you want depends on whether anyone is waiting for the answer. Sonnet against Opus covers the tier-choice question in more depth, and the Claude API pricing guide has the full rate card including batch.
Who should actually run Haiku 4.5
Run it for bounded, clearly specified code generation. Function-level generation, small refactors with a stated contract, test scaffolding, format conversion, one-shot fixes with a clear spec. On work shaped like that, it did everything the $4.05 model did for $0.94 and did it faster. Within the Claude family, on this evidence, there is no argument for starting anywhere else — and the burden of proof sits on the tier that wants 4.3x.
Do not run it, on this evidence, for the things the tiers exist for. Long-context work over a large codebase, sustained multi-step reasoning, agentic tool use across many turns, ambiguous specifications where the model has to ask the right question. Our harness does not test any of that, so we have nothing to say about it, and the fact that Haiku 4.5 tied on nine short functions is not evidence it will tie on a 200,000-token refactor. It is evidence that nine short functions do not separate these models.
The practical version: default to Haiku 4.5 for the bounded work, keep a bigger tier available for the rest, and measure the boundary on your own tasks rather than taking ours. Route by task type, not by reputation. If you want to run the same comparison on your own token mix, the cost calculator does the arithmetic.
Test the whole Claude ladder on one key
One OpenAI-compatible endpoint, 300+ models, one key. Swap the model id between Haiku 4.5, Sonnet 5 and Opus 4.8, run your own tasks, compare the bill.
How these numbers were produced
Nine executed Python tasks: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line.
The model receives a function signature and a prose spec. It never sees the assertions. The returned code runs against hidden assertions in an isolated python3 -I subprocess with a 12-second timeout. Every assertion passes or the task fails — no partial credit, no human grader, no LLM judge.
Temperature 0, max_tokens 4000. One scored attempt per task. The harness retries only on an API error, never on a wrong answer, which is why a miss would have stayed a miss.
Cost is derived, not invoiced. It is the token counts the API reported multiplied by that model's list price on a stated date. Haiku 4.5 ran on 2026-07-29 and is priced at 2026-07-29 rates. List prices move — 49 of roughly 396 catalogue models changed price in the twelve days to 2026-07-29 — so every figure here is true as of its pricing date and should be recomputed against current rates before you act on it.
Calls go through OpenRouter's OpenAI-compatible endpoint, deliberately not through the DataLLM Lab gateway. Nothing on this page depends on our infrastructure and you do not have to be our customer to reproduce it.
Our original sweep was 13 models run in one sitting. Haiku 4.5 is one of eight models run later on the same harness under the same settings, which brings the total to 21. Where this page says 21 models or 18 of 21, that is the combined set; the core sweep was and remains 13.
What we did not measure
Not measured at all: long-context reasoning, multi-file refactoring, agentic or multi-turn tool use, any non-Python language, and vision. The harness is single-turn. It does not run agents and it does not use tools. Everything the Claude tiers are differentiated on commercially sits in that list.
A known scoring artefact: the 4,000-token ceiling can truncate a very verbose model mid-answer, and a truncated answer scores as a miss. That penalises verbosity, which is a real production cost but is not the same thing as being wrong. Haiku 4.5 was not affected — it scored 9/9 — but the ceiling is part of the harness and you should know it is there.
Nine tasks is nine data points. A 9/9 is a clean result on this set, not a claim about a distribution. We did not run it twice, we did not vary the prompts, and we did not buy any model a retry.
Not tested, and never claimed as ours: GPT-5 nano, Ollama or any locally-run model, Grok 3, Grok 4, Grok 4.20, grok-build-0.1, VibeThinker, and any Gemini other than 3.6 Flash and 3.1 Pro. Claude Fable 5 and Claude Opus 4.7 have not been through this harness either. If a first-party number for any of those appears anywhere on this site, it is an error.
FAQ
Is Claude Haiku 4.5 good enough for coding?
On our benchmark it scored 9/9 on nine executed Python tasks — the same score as Claude Sonnet 5 and Claude Opus 4.8 — at $0.94 per 1,000 tasks priced 2026-07-29. For bounded, clearly specified function-level work, yes. For long-context refactoring, agentic tool use or ambiguous specs, we have no data either way, because this harness does not test those. Test it on your own workload before defaulting to a bigger tier.
How much cheaper is Haiku 4.5 than Opus 4.8?
4.3x on our measured cost: $0.94 against $4.05 per 1,000 tasks for the same nine tasks and the same 9/9 score. On list price the ratio is 5x — $1 / $5 per 1M against $5 / $25, both captured 2026-07-29. Both models emitted zero reasoning tokens, so neither inflated its own bill with a thinking trace; the measured gap lands under the sticker gap because Opus 4.8 returned somewhat fewer tokens across the nine tasks.
Why did the expensive Claude models not score higher?
Because the tasks are not hard enough to separate them. Nine short, self-contained Python functions are inside the competent range of every serious coding model shipping in 2026 — 18 of the 21 models we have run scored 9/9. The Claude tiers are differentiated on long context, sustained reasoning, agentic tool use and ambiguous specs, and this harness exercises none of those. A tie here means the test did not reach them, not that the tiers are identical.
What does zero reasoning tokens mean in practice?
Haiku 4.5 answered directly on all nine tasks without emitting a reasoning trace. Reasoning tokens bill at the output rate and take wall-clock time to produce, so a model at zero is both cheaper and faster, and — the part people miss — it has no cost variance from prompt to prompt. For comparison, Gemini 3.1 Pro averaged 1,094 reasoning tokens on the same nine tasks and cost $14.70 per 1,000 tasks priced 2026-07-29, a 15.6x gap for the same 9/9.
Should I move from Claude Sonnet 4.6 to Sonnet 5?
On this workload the newer model is cheaper: Sonnet 5 measured $1.67 per 1,000 tasks against Sonnet 4.6 at $2.22, both at 9/9, because Sonnet 5 lists at $2 / $10 per 1M and Sonnet 4.6 at $3 / $15 as captured 2026-07-29. The one thing you give up is latency — Sonnet 4.6 averaged 4.9 s and Sonnet 5 averaged 7.2 s. Cheaper and equally correct, slightly slower.
Is this measured cost the same as my bill?
No. We take the token counts the API reported and multiply by each model's list price on a stated date — 2026-07-29 for Haiku 4.5. It is a derived measured cost, not a vendor invoice. Your bill will differ with prompt length, caching, discounts, provider routing and any price change since that date. The ratios between models are the durable part; the absolute dollars are not.
DataLLM Lab