Model Comparison

Claude vs GPT-5 in 2026: We Ran Nine Tiers Across Both Families (All 9/9, $0.94 to $8.83)

Claude and GPT-5 are the two leading frontier families of 2026, and most comparisons end in a verdict nobody can check. So we ran them. Five Claude tiers and four GPT-5 tiers went through the same executed nine-task Python benchmark in July 2026, and all nine scored 9/9. On work that size the family choice is not a capability choice at all — it is a cost and latency choice, and there the spread is $0.94 to $8.83 per 1,000 tasks. That result is narrow on purpose: nine short, self-contained functions say nothing about long context, multi-file refactoring or agentic tool use, which is exactly where the two families genuinely diverge. This guide gives you our measured numbers, the independent third-party benchmarks, modeled monthly costs, and worked scenarios for picking by job.

Claude vs GPT-5 — benchmarks, measured cost per 1,000 tasks, and which to use by job

The short answer

On short, well-specified coding work, both families are the same model. We ran nine tiers across Claude and GPT-5 on the same executed benchmark and every one returned 9/9. What separates them there is money: Claude Haiku 4.5 at $0.94 and GPT-5.4 at $1.69 per 1,000 tasks, against Claude Opus 5 at $5.64 and GPT-5.5 at $8.83 for the identical result.

Where the families genuinely diverge is everything our harness does not reach — long context, multi-file refactoring, agentic tool use, non-Python work, writing quality. There the third-party picture still splits: independent SWE-bench Verified puts Claude Opus ahead on repository bug-fixing, Terminal-Bench puts GPT-5 ahead on shell-driving agents. For most teams the real answer stays both, routed by task — with the cheap tier doing far more of the work than its price suggests.

How this is sourced. The 9/9 scores, latencies, token counts and measured costs are our own executed runs — method in how our numbers were produced and on the methodology page. SWE-bench Verified and Terminal-Bench figures are third-party and attributed inline. List prices are from the DataLLM Lab catalog captured 2026-07-29, except Opus 5, priced from the same catalog on 2026-07-30. Deeper dives: Claude API guide, GPT-5 API guide.

What we measured: nine tiers, one score

Across July 2026, five Claude tiers and four GPT-5 tiers went through the same executed nine-task Python harness. The model gets a function signature and a prose spec, never sees the assertions, and the returned code is executed against them. One scored attempt each, temperature 0.

All nine scored 9/9. No misses, in either family, at any tier. Rows are ordered by measured cost per 1,000 tasks, cheapest first. The last column is the date of the list price each cost was computed at; Sonnet 5, Opus 4.8 and GPT-5.5 came from our original 13-model sweep, the other six ran later on the same harness on the dates shown.

ModelFamilyScoreMeasured cost / 1k tasksMean latencyReasoning tokens per taskPriced at
Claude Haiku 4.5Claude9/9$0.943.7 s02026-07-29
GPT-5 miniGPT-59/9$1.5315.2 s5552026-07-29
Claude Sonnet 5Claude9/9$1.677.2 s02026-07-17
GPT-5.4GPT-59/9$1.693.6 s02026-07-29
Claude Sonnet 4.6Claude9/9$2.224.9 s02026-07-29
Claude Opus 4.8Claude9/9$4.056.1 s02026-07-17
GPT-5.6 SolGPT-59/9$4.986.6 s582026-07-29
Claude Opus 5Claude9/9$5.645.3 s62026-07-30
GPT-5.5GPT-59/9$8.8310.5 s1762026-07-17
Both families, same 9/9: measured cost per 1,000 tasksNine executed Python tasks, temperature 0, one scored attempt each. Blue = Claude, grey = GPT-5. Every bar scored 9/9.Claude Haiku 4.5$0.94GPT-5 mini$1.53Claude Sonnet 5$1.67GPT-5.4$1.69Claude Sonnet 4.6$2.22Claude Opus 4.8$4.05GPT-5.6 Sol$4.98Claude Opus 5$5.64GPT-5.5$8.83One scale throughout: 60 px per dollar. Cost is measured token counts multiplied by list price on the date in the table above, not a vendor invoice.
Chart: DataLLM Lab. Scores, latencies and token counts are measured on our executed nine-task Python benchmark; cost is those token counts multiplied by each model's list price on the date shown. Method: our methodology. Full run: the coding cost benchmark.

The score column never moves and the cost column moves 9.4x. That is the finding. The two flagships people actually reach for cost 6.0x and 9.4x what Haiku 4.5 cost for the same nine results — Claude Opus 5 at $5.64 and GPT-5.5 at $8.83 against $0.94.

Three things in that table are invisible on any rate card, because they come from token counts rather than prices:

Cross-family, at both ends of the ladder, Claude was the cheaper option on this workload: $0.94 against $1.53 at the cheap tier, $5.64 against $8.83 at the flagship. On latency GPT-5.4 was the quickest of the nine at 3.6 s, one tenth of a second ahead of Claude Haiku 4.5, and GPT-5 mini was the slowest at 15.2 s. Latency and cost do not rank the same way, which is the same disagreement we found across the whole field in the AI coding ranking.

These nine are part of a larger set. Our original sweep was 13 models run in one sitting, of which 10 scored 9/9; ten more have since run on the same harness under the same settings, and 20 of the 23 scored 9/9. Across all 23 the measured spread is $0.10 to $14.70. Full run: the LLM coding cost benchmark.

What nine short functions cannot tell you

Read the section above with its scope attached, because the scope is small. The harness runs nine short, self-contained Python functions. Every task fits in a few hundred tokens, is single-turn, and has an unambiguous spec.

It does not measure: long-context work, multi-file refactoring, agentic or multi-turn tool use, any non-Python language, vision, translation, or writing quality. It does not run agents and it does not use tools.

That list is not a footnote — it is the entire commercial argument for the expensive tiers, and it is where the two families actually differ. Anthropic sells Opus for sustained multi-step reasoning and long context; OpenAI sells GPT-5.5 and the Codex variants for agentic execution. Our nine tasks sit far below the level at which either claim starts to be tested. A tie here means the test did not reach them, not that the tiers are interchangeable.

So do not read this page as "just use Haiku". Read it as: for bounded, clearly specified code generation, the cheap tier of either family did everything the flagship did, and the burden of proof sits on the tier that wants 6x to 9.4x. Find where your own workload stops being easy — that is the only place the tier question has an answer, and it is not a place a nine-task benchmark can find for you.

Side by side

Claude (Anthropic)GPT-5 (OpenAI)
FlagshipOpus 5 / Opus 4.8GPT-5.6 Sol / GPT-5.5
Our nine-task score, every tier run9/99/9
Our measured cost, cheap tier$0.94 (Haiku 4.5)$1.53 (GPT-5 mini)
Our measured cost, flagship$5.64 (Opus 5)$8.83 (GPT-5.5)
Our fastest tier measured3.7 s (Haiku 4.5)3.6 s (GPT-5.4)
SWE-bench Verified (third-party)88.6% (Opus 4.8)82.6% (GPT-5.5)
Terminal-Bench (third-party)StrongLeads (~82.7%)
Flagship list price (in/out)$5 / $25$5 / $30 (5.5, 5.6 Sol) · $2.50 / $15 (5.4)
Cheap tier list priceHaiku 4.5 $1 / $5GPT-5 mini $0.25 / $2
Strength (not measured by us)Long context, planning, instruction-followingTerminal agents, tool use, ecosystem

Note the row that flips: GPT-5 mini has the cheaper list price — 4x on input and 2.5x on output — and Claude Haiku 4.5 has the cheaper measured cost by 1.6x. Reasoning tokens are the whole difference. Per-token rate cards cannot show you that, which is the reason we run the tasks.

Third-party benchmarks

Our harness answers a narrow question. For the broader one, the independent boards split. On the independent SWE-bench Verified board, Claude Opus 4.8 (88.6%) leads GPT-5.5 (82.6%) on real-world repository bug-fixing:

Claude vs GPT-5 — independent SWE-bench Verifiedvals.ai, June 2026 — third-party, not our runClaude Opus 4.888.6%GPT-5.582.6%
Chart: DataLLM Lab — data from independent SWE-bench Verified (vals.ai), June 2026. Not our measurement. Claude Opus 4.8 leads on real-world coding; GPT-5 tends to lead separate agentic terminal benchmarks (Terminal-Bench).

The nuance behind the headline: SWE-bench measures bug-fixing against real repositories, where Claude leads; Terminal-Bench measures command-line agent execution — running commands, navigating a shell, recovering from errors — where GPT-5 leads (~82.7%). So "better at coding" genuinely depends on whether your work is patching code or driving a terminal. Treat a single benchmark as one data point, not the verdict, and weight the one that matches your workload.

Read against our run, the two sources are consistent rather than contradictory. SWE-bench tasks are large, multi-file and ambiguous; ours are small, single-file and precise. The families separate on the first kind of work and tie on the second. That is the practical shape of the whole comparison: the harder the task, the more the family choice matters — and the more of your traffic is easy, the less you should be paying for it. Rankings across the field are compared in best coding LLM 2026.

What they cost to run

List prices are close enough at the flagship tier that they rarely decide it — but the tier you pick within each family changes the bill by an order of magnitude. Here is the modeled monthly cost across five workloads, from list price and stated token volumes:

Monthly workloadClaude Opus 5Claude Sonnet 4.6GPT-5.5GPT-5.4GPT-5 mini
Support chatbot$500$300$560$280$34.0
RAG / knowledge base$1,500$900$1,600$800$90.0
Coding agent$1,025$615$1,150$575$70.0
Batch extraction$950$570$990$495$53.5
Content generation$1,100$660$1,300$650$85.0
Methodology. Modeled, not measured. Cost = input_price × input volume + output_price × output volume, at list prices captured 2026-07-29. Monthly volumes: Support chatbot 40M in / 12M out, RAG 200M / 20M, Coding agent 80M / 25M, Batch extraction 150M / 8M, Content generation 20M / 40M. Claude Opus 4.8 lists identically to Opus 5 ($5 / $25) so its column would be the same — which is exactly the gap our measured run exposes.

Opus 5 and GPT-5.5 land within 4% to 18% of each other across these five workloads — close enough that list price rarely decides the flagship; GPT-5.4 is the cheapest frontier option on list price. But the real lever is dropping to a cheap tier where quality allows — GPT-5 mini models a chatbot at $34 against Opus at $500.

Two cautions on that table. First, it assumes a fixed token volume, and our measured run shows models with identical list prices consuming very different token counts for the same work — Opus 5 versus Opus 4.8 differed by 39%, GPT-5.5 versus GPT-5.6 Sol by 1.8x. Second, reasoning tokens are output tokens: GPT-5 mini's 555 per task on our harness is why its modeled $34 will understate a reasoning-heavy workload. Run your own mix through the cost calculator rather than trusting either table as a quote. Prices also move — 49 of roughly 396 catalogue models changed price in the twelve days to 2026-07-29.

Where Claude wins

Where GPT-5 wins

Worked scenarios

Scenario Bounded code generation

  • Function-level work with a clear spec → Claude Haiku 4.5 ($0.94) or GPT-5.4 ($1.69). Both scored 9/9 on our harness; the flagships scored the same for 6.0x to 9.4x Haiku's measured cost.

Scenario Production coding agent

  • Long context and ambiguity → Claude Opus for the hard reasoning (SWE-bench lead), a cheap tier for routine reads. GPT-5 Codex if the loop is terminal-heavy. Our harness does not test this.

Scenario High-volume chatbot

  • Cost dominates → Claude Haiku 4.5 or GPT-5 mini. Check reasoning tokens, not just list price: mini spent 555 per task on our nine and 15.2 s per call.

Scenario Mixed product

  • Route both — cheap tier for the bulk, Claude for long-context code, GPT-5 for terminal and tools. One key, per-request choice.

Which to pick by job

Coding

  • Short specified functions: either cheap tier, measured identical. Large refactors: Claude Opus. Terminal loops: GPT-5 Codex. See best coding LLM.

Agents

  • GPT-5 for terminal/tool execution; Claude for planning and computer-use. We measured neither. See best LLM for agents.

Writing

  • Both excellent — preference is subjective and our benchmark says nothing about it. Use the cheaper tiers and test on your voice. See best LLM for writing.

Best move Route by difficulty

  • Send the easy majority to a cheap tier, escalate the rest. The tier question inside Claude is worked through in Sonnet vs Opus.

Run Claude and GPT-5 side by side

Claude Opus 5, Claude Haiku 4.5, GPT-5.5, GPT-5.4 and 300+ more — one OpenAI-compatible key, live price comparison, route per request.

How our numbers were produced

Nine executed Python tasks: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line.

The model receives a function signature and a prose spec. It never sees the assertions. The returned code runs against hidden assertions in an isolated python3 -I subprocess with a 12-second timeout. Every assertion passes or the task fails — no partial credit, no human grader, no LLM judge. Temperature 0, max_tokens 4000, one scored attempt per task; the harness retries only on an API error, never on a wrong answer.

Cost is derived, not invoiced. It is the token counts the API reported multiplied by that model's list price on the date shown in the table — 2026-07-17 for Sonnet 5, Opus 4.8 and GPT-5.5; 2026-07-29 for Haiku 4.5, Sonnet 4.6, GPT-5 mini, GPT-5.4 and GPT-5.6 Sol; 2026-07-30 for Opus 5. Write it down as measured cost, not a bill. Prices move, so recompute against current rates before acting on any figure here.

Calls go through OpenRouter's OpenAI-compatible endpoint, deliberately not through the DataLLM Lab gateway. Nothing on this page depends on our infrastructure and you do not have to be our customer to reproduce it.

Our original sweep was 13 models run in one sitting, 10 of which scored 9/9. Ten more have since run on the same harness under the same settings, bringing the total to 23, of which 20 scored 9/9. The core sweep was and remains 13.

Not tested, and never claimed as ours: Claude Fable 5, Claude Opus 4.7, GPT-5 nano, the Codex variants, Ollama or any locally-run model, Grok 3, Grok 4, Grok 4.20, grok-build-0.1, VibeThinker, and any Gemini other than 3.6 Flash and 3.1 Pro. One known artefact: the 4,000-token ceiling can truncate a very verbose model mid-answer and a truncated answer scores as a miss. None of the nine models on this page were affected.

FAQ

Is Claude better than GPT-5?

On our own executed benchmark, no — and neither is GPT-5 better than Claude. Five Claude tiers and four GPT-5 tiers all scored 9/9 on the same nine Python tasks. On larger, more ambiguous work the third-party picture splits: Claude Opus 4.8 (88.6%) leads GPT-5.5 (82.6%) on independent SWE-bench Verified, while GPT-5 leads Terminal-Bench for shell-driving agents. Short specified coding: pick on cost. Large refactors: lean Claude. Terminal agents: lean GPT-5.

Is Claude or GPT-5 cheaper?

On our measured runs Claude was cheaper at both ends of the ladder: Haiku 4.5 $0.94 against GPT-5 mini $1.53 per 1,000 tasks, and Opus 5 $5.64 against GPT-5.5 $8.83. On list price it goes the other way at the cheap tier — GPT-5 mini is $0.25 / $2 per 1M against Haiku 4.5 at $1 / $5, captured 2026-07-29. The difference is reasoning tokens, which bill at the output rate: mini spent 555 per task and Haiku 4.5 spent zero.

Claude or GPT-5 for coding?

Depends on task size. For bounded, clearly specified functions our nine-task run cannot separate them, so take the cheap tier of either family — Claude Haiku 4.5 at $0.94 or GPT-5.4 at $1.69 per 1,000 tasks, both 9/9. For large multi-file work, Claude Opus leads independent SWE-bench Verified. For terminal-heavy agent loops, test GPT-5 Codex, which we have not run.

Did every tier really score the same?

Yes, on this harness: all nine returned 9/9. Across all 23 models we have run, 20 scored 9/9. That says more about the tasks than the models — nine short, self-contained Python functions sit inside the competent range of every serious 2026 coding model. The tiers are differentiated on long context, sustained reasoning, agentic tool use and ambiguous specs, and this harness exercises none of those. A tie means the test did not reach them.

Claude or GPT-5 for agents and writing?

We have no first-party data on either. Our harness is single-turn, Python-only, and does not use tools, so it cannot speak to agentic execution or writing quality. Third-party: GPT-5 (5.5 and Codex) tends to lead terminal and tool-use benchmarks; Claude Opus leads planning and computer-use. On writing both are strong and preference is subjective — test both on your own voice with the cheaper tiers first.

Can I use both Claude and GPT-5 with one API?

Yes — through an OpenAI-compatible gateway like DataLLM Lab you reach Claude Opus 5, Claude Haiku 4.5, GPT-5.5, GPT-5.4 and 300+ others with one key, so you can route by task and by difficulty without maintaining two integrations. Note that our benchmark runs deliberately go through OpenRouter, not through our own gateway, so none of the numbers above depend on it.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.