Model Comparison

GPT-5.6 Pro Measured: Luna Pro 8/9 at $3.79, Terra Pro 9/9 at $6.99 (and We Did Not Run Sol Pro)

The GPT-5.6 Pro tier ships in three flavours — Luna Pro, Terra Pro and Sol Pro, all three listed in our 2026-07-29 catalogue capture — and on 2026-08-06 we ran two of them through our executed nine-task Python benchmark. GPT-5.6 Luna Pro scored 8/9 at a measured $3.79 per 1,000 tasks, 7.2 s average, 548 reasoning tokens per task. GPT-5.6 Terra Pro scored 9/9 at $6.99, 5.0 s, 246 reasoning tokens. We did not run Sol Pro, so this page says nothing about it. The result worth your attention is the inversion: the cheaper Pro model is the one that missed a task, and the plain non-Pro GPT-5.6 Sol sat between them at $4.98 with a clean 9/9 and only 58 reasoning tokens — the least thinking and the cheapest perfect score of the three.

Bar chart of measured cost per 1,000 tasks for GPT-5.4, GPT-5.6 Luna Pro, GPT-5.6 Sol and GPT-5.6 Terra Pro

This site had no first-party coverage of the GPT-5.6 Pro tier until now. On 2026-08-06 we put two of the three through the same executed harness every other model on this site runs on. Two of three, not three of three — that gap is stated up front and repeated below, because a page that quietly extrapolates from two models to a third is worth nothing.

The headline is an inversion. Inside one vendor family, on this workload, the Pro badge did not buy a better score, and the cheapest of the three Pro-or-not models we measured is the one that missed a task. Everything below is arithmetic on numbers we recorded.

What we ran on 2026-08-06

Two GPT-5.6 Pro models, both priced at rates captured from the live catalogue:

For reference we had already run the non-Pro openai/gpt-5.6-sol on 2026-07-29: 9/9, $4.98 per 1,000 tasks, 6.6 s, 58 reasoning tokens per task, priced 2026-07-29 at $5 / $30 per 1M. It is a different tier and a different run date, and we flag both every time it appears.

We did not run Sol Pro. Not on 2026-08-06, not since. There is no Sol Pro row in our benchmark file and nothing on this page infers one.

The limit, stated before the numbers

Read this before the tables. Our harness is nine short, self-contained Python functions. One signature, one prose spec, one scored attempt, hidden assertions, temperature 0. A Pro tier exists for long-context work, hard multi-step reasoning and agentic loops — none of which this benchmark touches. A single missed spec-following task is thin evidence about a model's ceiling, and we are not going to pretend otherwise. What follows is a precise result on a narrow workload, which is a different thing from a verdict on a model.

That caveat cuts in the direction you would expect: it means Luna Pro's 8/9 does not prove Luna Pro is weak, and Terra Pro's 9/9 does not prove Terra Pro is strong. What it does prove is that on bounded generate-and-test coding, paying for the Pro rung of this family did not produce a better score than not paying for it — because the non-Pro Sol matched Terra Pro at 9/9 and beat Luna Pro outright.

The cheaper Pro model scored worse

Luna Pro cost $3.79 per 1,000 tasks and missed a task. Terra Pro cost $6.99 and did not. That is the ordinary shape of a price ladder, and it is the last ordinary thing on this page. Sol, which carries no Pro badge at all, landed between them on cost at $4.98, scored a clean 9/9, and did it on 58 reasoning tokens per task — a fraction of what either Pro model spent thinking.

ModelScoreList price in / out per 1MMeasured cost / 1k tasksMean latencyReasoning tokens / taskMissedPriced at
GPT-5.6 Luna Pro8/9$0.50 / $3$3.797.2 s548parse_csv_line2026-08-06
GPT-5.6 Sol · not Pro9/9$5 / $30$4.986.6 s58none2026-07-29
GPT-5.6 Terra Pro9/9$1.25 / $7.50$6.995.0 s246none2026-08-06
GPT-5.6 Sol Pronot run$5 / $30not run2026-07-29 rate only

The fourth row is not padding. Sol Pro is a real model we have not measured, and leaving it out of the table would let a reader assume the tier had been swept. It has not been.

About the two pricing dates, because a stale price would be a boring explanation for an interesting gap. Our catalogue capture of 2026-07-29 lists Luna Pro at $0.50 / $3 and Terra Pro at $1.25 / $7.50byte-identical to the rates recorded against the 2026-08-06 runs. So for the two Pro models nothing moved in that eight-day window, and the cost comparison is clean. Sol's $5 / $30 comes from that same 2026-07-29 capture, which is also its run date; we have no later capture of Sol on disk, so we will say only that all three figures rest on the same catalogue snapshot, two of them re-confirmed eight days on. Prices move, and we argue for re-checking them in LLM price volatility.

One more fact from that capture, and it is the one that reframes the whole tier: in our 2026-07-29 catalogue snapshot each Pro id lists at exactly the same per-token rate as its non-Pro sibling. Luna and Luna Pro both $0.50 / $3. Terra and Terra Pro both $1.25 / $7.50. Sol and Sol Pro both $5 / $30. Pro costs nothing extra per token. Whatever a Pro id does to your bill therefore has to arrive as token volume rather than as rate — and we cannot show you that happening, because we ran no Pro/non-Pro sibling pair. We have Luna Pro but not Luna, Terra Pro but not Terra, Sol but not Sol Pro. What we can show is that a rate card this flat tells you nothing about what a model will spend, which is why the only unit worth ranking on is measured cost per task actually run, read next to the score.

Where the rate card stops predicting the bill

Hold Sol at 1.00x and look at what the rate card promises against what we measured.

ModelScoreList rate vs SolMeasured cost / 1k tasksMeasured vs SolGap between the two
GPT-5.6 Luna Pro8/90.10x$3.790.76x7.6x worse than promised
GPT-5.6 Sol · baseline9/91.00x$4.981.00x
GPT-5.6 Terra Pro9/90.25x$6.991.40x5.6x worse than promised

Luna Pro lists at one tenth of Sol's per-token rate and cost 76% of Sol per 1,000 tasks run. A 10x discount on the rate card collapsed to a 24% discount in practice — and bought a lower score. Charge it for that miss and the discount shrinks again: Luna Pro passed 8 of 9, so its $3.79 works out to $4.26 per 1,000 tasks it actually got right against Sol's $4.98, a 14% saving rather than 24%. It is still the cheaper of the two either way. Terra Pro lists at one quarter of Sol's rate and cost 1.40x as much per completed task — both scored 9/9, so for those two the two units are the same number. A 4x discount on paper turned into a 40% premium in reality, at an identical score.

There is no mystery in the mechanism, only in the fact that it is invisible until you run it. Cost is tokens multiplied by rate. Rate went down. Cost went up. Therefore token count went up, by more than the rate went down. We hold that logic to the arithmetic and no further: our run record does not store a token split for the Sol run, so we will not back one out and print a reconstruction as a measurement.

What we can show directly is the split for the two Pro models, and it is where the story lands.

Where the money actually went

Across the nine tasks the API reported 21,729 input and 7,748 output tokens for Luna Pro, against 19,726 input and 5,102 output for Terra Pro. Reasoning tokens are counted inside that output total and bill at the output rate, which is the expensive one.

ModelScoreInput tokens / taskOutput tokens / taskReasoning tokens / taskReasoning as share of outputReasoning cost / 1k tasks
GPT-5.6 Luna Pro8/92,41486154864%$1.64
GPT-5.6 Terra Pro9/92,19256724643%$1.85
GPT-5.6 Sol · not Pro9/9not storednot stored58$1.74

Luna Pro returned 52% more output tokens than Terra Pro while scoring one task lower — 861 per task against 567 — and 64% of everything it emitted was reasoning it billed for and did not ship. Terra Pro spent 43% of its output on reasoning. Sol spent 58 tokens a task and scored 9/9.

The last column is the part worth sitting with. Multiply reasoning tokens per task by 1,000 tasks and by each model's output rate and you get what the thinking alone cost: $1.64 for Luna Pro, $1.85 for Terra Pro, $1.74 for Sol. Three models, reasoning budgets ranging from 58 tokens to 548 — a 9.4x spread — and a reasoning bill that lands inside a 13% band on all three, because the cheap models think a lot at a cheap rate and the expensive one barely thinks at an expensive rate. Sol's frugality is doing the same job Luna Pro's discount is doing, and Sol got the better score for it.

If you want this arithmetic on your own token mix rather than ours, the cost calculator does it.

Terra Pro against the wider field

Now widen out past one vendor family. GPT-5.4 scored the same 9/9 as Terra Pro at a measured $1.69 per 1,000 tasks against $6.99 — 4.1x less, priced 2026-07-29 and 2026-08-06 respectively. It was also faster: 3.6 s against 5.0 s, and it emitted zero reasoning tokens.

The rate card again points the wrong way. Terra Pro lists at $1.25 / $7.50 — exactly half GPT-5.4's $2.50 / $15 on both input and output — and cost 4.1x more per completed task. That is an 8.3x swing between what the per-token price implies and what we measured, at an identical score. If you have been tracking the GPT-5 line, this is the same shape as the generational question in GPT-5.4 vs GPT-5.5: the newer, nominally cheaper thing is not reliably the cheaper thing.

The Pro tier cost more and did not score better: $1.69 to $6.99 per 1,000 tasksNine executed Python tasks, temperature 0, one scored attempt each. Blue bars are the two GPT-5.6 Pro models we ran.GPT-5.4 · 9/9$1.69GPT-5.6 Luna Pro · 8/9$3.79GPT-5.6 Sol · 9/9$4.98GPT-5.6 Terra Pro · 9/9$6.99One scale throughout: 70 px per dollar. Cost is measured token counts multiplied by list price on the date stated in the tables above.GPT-5.4 and GPT-5.6 Sol priced 2026-07-29; Luna Pro and Terra Pro priced 2026-08-06. GPT-5.6 Sol Pro is absent because we did not run it.
Chart: DataLLM Lab. Scores, latencies and token counts are measured on our executed nine-task benchmark; cost is those token counts multiplied by each model's list price on the stated date, not a vendor invoice. Method: our methodology. Full run: the coding cost benchmark.

Placed in the whole set: as of 2026-08-06 we had 45 usable model results, of which 33 scored 9/9. Terra Pro's $6.99 is the 8th highest measured cost of those 45, and 27 of the 33 models that also scored 9/9 cost less than it did. Only five 9/9 models cost more. Luna Pro is harsher still: 21 of the 33 perfect scorers came in under its $3.79, and it did not score 9/9. Its 7.2 s is the slowest of the three GPT-5.6 models in our file. The cheapest 9/9 we have recorded is DeepSeek V3.2 at $0.08 priced 2026-07-30, and the fastest 9/9 in that set is GPT-5.4 mini at 2.3 s. The full sort by every axis is in the AI coding ranking.

None of that makes Terra Pro a bad model. It makes nine bounded Python functions a workload where 33 of 45 models are already at the ceiling, so price and latency are the only columns still moving. Ranking a Pro tier on a test it cannot fail differently is not a capability verdict.

Sol Pro: we did not run it

There is one more model in this tier and we have no data on it. openai/gpt-5.6-sol-pro has never been through our harness. It has no row in our benchmark file, no score, no measured cost, no latency and no reasoning token count.

The temptation here is obvious: we measured two thirds of a tier, both Pro models behaved a particular way, so surely the third does too. We are not doing that, and the two models we did run are the reason why. Luna Pro and Terra Pro are siblings on the same ladder released together, and they diverged on score, on latency, on reasoning budget and on the direction of their cost error against the rate card. Two members of this family predicted each other badly. Extrapolating to a third from that base would be worse than saying nothing.

What we can state without measuring: our 2026-07-29 catalogue capture lists Sol Pro at $5 / $30 per 1M, identical to non-Pro Sol. That is a rate, not a result. Everything else about Sol Pro on this page is a blank, and it will stay a blank until it runs. For what OpenAI itself says about Sol, Terra and Luna — the non-Pro tiers — our GPT-5.6 family guide collects the vendor claims with dates attached. It does not cover the Pro ids either; nothing on this site does beyond the two rows above.

What the Pro tier bought on this workload

On these nine tasks, Pro bought latency and price, not correctness. Terra Pro was the fastest of the three at 5.0 s and the most expensive at $6.99. Luna Pro was the cheapest at $3.79 and the only one to miss. Sol, with no Pro badge, was in the middle on cost and tied for the top on score. That is a familiar shape on this site: inside a single vendor's ladder, a higher tier on bounded work reliably moves the cost and latency columns and rarely moves the score column.

If your workload looks like ours — short, well-specified, single-turn code generation — do not buy the Pro rung. GPT-5.4 delivered the same 9/9 at $1.69 and 3.6 s. Twenty-one models in our set delivered 9/9 for less than Luna Pro charged to deliver 8/9. The cheap coding model roundup is the more useful page for that tier of work.

If your workload does not look like ours, this page does not answer your question and should not pretend to. Long context, multi-file refactors, agent loops running for an hour — that is what a Pro tier is sold for, and our harness never goes near it. The honest instruction is to run your own prompts and count your own tokens, because the one thing we did demonstrate is that the rate card gets the ordering wrong.

The general rule. Rank by measured cost per 1,000 tasks run, read next to the score, not by per-token price, and never by tier name. Inside GPT-5.6 the Pro ids carry the same per-token rate as their non-Pro siblings, so nothing on a pricing page separates a Pro id from its twin — any difference has to live in token counts no vendor publishes, and we did not run a sibling pair to measure one. Luna Pro and Terra Pro differed by 1.8x on the same nine tasks, and the cheaper one is the one that scored worse.

Count your own reasoning tokens

Every figure above is a token count multiplied by a list price, and you can produce the same numbers for your own prompts on whichever endpoint carries these ids. For the models you route every day, DataLLM Lab puts 300+ of them behind one OpenAI-compatible key at list prices.

How these numbers were produced

Nine executed Python tasks: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line.

The model receives a function signature and a prose spec. It never sees the assertions. Returned code runs against hidden assertions in an isolated python3 -I subprocess with a 12-second timeout. Every assertion passes or the task fails — no partial credit, no human grader, no LLM judge. Temperature 0, max_tokens 4000, one scored attempt per task, retrying only on an API error and never on a wrong answer. That is why Luna Pro's miss stayed a miss.

The task Luna Pro missed is the one that separates this field. parse_csv_line was missed by 9 of the 45 usable models — quoted fields, embedded commas, escaped quotes — and it is the single most-missed task in the set. Luna Pro joins Llama 4 Scout, gpt-oss-120b, Gemini 3 Flash Preview, MiMo v2.5, MiniMax M2.5, DeepSeek V4-Pro, Gemini 3.5 Flash and Gemini 2.5 Pro on it. So the miss is not random noise on an easy item; it is the hardest item in the set. It is still one item.

Cost is derived, not invoiced. It is the token counts the API reported multiplied by that model's list price on a stated date. Luna Pro and Terra Pro ran 2026-08-06 at 2026-08-06 rates; GPT-5.6 Sol and GPT-5.4 ran 2026-07-29 at 2026-07-29 rates. We write measured cost, never billed cost, and every figure should be recomputed before you act on it.

Calls go through OpenRouter's OpenAI-compatible endpoint, deliberately not through the DataLLM Lab gateway. Nothing on this page depends on our infrastructure and you do not have to be our customer to reproduce it.

On the counting. Our core sweep was 13 models run in one sitting, and that number stays what it was. Everything since, including both GPT-5.6 Pro models, ran later on the same harness under the same settings. As of 2026-08-06 the file holds 47 entries: 45 usable and 2 excluded. Where this page says 45 models or 33 of 45, that is the combined set.

The two exclusions, and why they are not scores. Claude Fable 5 returned 4/5 because four tasks came back empty with finish_reason=content_filter after three retries — a harness artefact we wrote up in the content-filter bug. Qwen3.5-397B returned 8/8 because token_bucket came back empty with finish_reason=length: the model spent our entire 4,000-token ceiling on reasoning without emitting an answer. That ceiling is our constraint, not a defect in the model. Neither partial result is comparable to a 9/9 and neither appears in any table above.

What we did not measure

Sol Pro, at all. Repeated here because it is the most likely thing to be misread off this page. No run, no score, no cost.

Not measured for any model: long-context behaviour, multi-file refactoring, agentic or multi-turn tool use, any non-Python language, vision, and anything run locally. The harness is single-turn and API-hosted. The capabilities a Pro tier is sold for sit entirely inside that list.

Reasoning effort is untested. We sent plain requests with no effort or thinking-budget parameter. Luna Pro's 548 and Terra Pro's 246 reasoning tokens per task are what each model chose to spend on nine bounded function specs at temperature 0. They tell you nothing about cost or quality at a higher effort setting, and since reasoning bills at the output rate, our cost figures should be read as a floor for these models rather than a typical figure.

Nine tasks is nine data points. An 8/9 is one miss on one item, not a defect rate. We did not run it twice, we did not vary the prompts, and we did not buy any model a retry on a wrong answer.

The 4,000-token ceiling is part of the harness. A model verbose enough to be truncated mid-answer scores that task as a miss. Neither Pro model was truncated — Luna Pro's miss was a wrong answer, not an empty one — but the ceiling penalises verbosity, and verbosity is exactly what this page is about.

FAQ

Did you benchmark GPT-5.6 Sol Pro?

No. We ran GPT-5.6 Luna Pro and GPT-5.6 Terra Pro on 2026-08-06 and that is all. Sol Pro has never been through our harness, has no row in our benchmark file, and we will not infer its behaviour from the two models we did run — those two diverged from each other on score, latency, reasoning budget and cost direction, so they are poor predictors of a sibling. Our 2026-07-29 catalogue capture lists Sol Pro at $5 / $30 per 1M, which is a rate, not a result.

Is GPT-5.6 Luna Pro cheaper than GPT-5.6 Terra Pro?

Yes, on both the rate card and our measurement — but it scored lower. Luna Pro lists at $0.50 / $3 per 1M and measured $3.79 per 1,000 tasks with an 8/9, missing parse_csv_line. Terra Pro lists at $1.25 / $7.50 and measured $6.99 with a 9/9. Both priced 2026-08-06. Luna Pro was also slower, 7.2 s against 5.0 s, and spent 548 reasoning tokens per task against Terra Pro's 246.

Is the GPT-5.6 Pro tier worth it for coding?

Not on the workload we measure. Nine short, self-contained Python functions: GPT-5.4 scored 9/9 at a measured $1.69 per 1,000 tasks and 3.6 s, against Terra Pro's 9/9 at $6.99 and 5.0 s — 4.1x less money for the same score. Of the 45 usable results we had as of 2026-08-06, 33 scored 9/9 and 27 of those cost less than Terra Pro. But a Pro tier is sold for long-context and hard multi-step work, and our harness never touches that, so this is a narrow answer to a narrow question.

Why did a model with a 10x cheaper rate card only cost 24% less?

Reasoning tokens. Luna Pro lists at one tenth of GPT-5.6 Sol's per-token rate yet measured $3.79 against Sol's $4.98 — 0.76x, not 0.10x. Luna Pro returned 861 output tokens per task and 548 of them, 64%, were reasoning it billed for and did not ship. Reasoning bills at the output rate. Multiply it out and the thinking alone cost $1.64 per 1,000 tasks for Luna Pro, $1.85 for Terra Pro and $1.74 for Sol — three models with reasoning budgets 9.4x apart landing within 13% of each other on the reasoning bill.

Does Pro cost more per token than non-Pro?

Not in our 2026-07-29 catalogue capture. Each Pro id listed at exactly the same rate as its non-Pro sibling: Luna and Luna Pro at $0.50 / $3, Terra and Terra Pro at $1.25 / $7.50, Sol and Sol Pro at $5 / $30, per 1M tokens. So nothing on the rate card separates a Pro id from its twin. Whether Pro costs more in practice would have to show up as token volume, and we cannot say that it does: we never ran a Pro/non-Pro sibling pair — we have Luna Pro but not Luna, Terra Pro but not Terra, Sol but not Sol Pro. It is also why measured cost per 1,000 tasks, read next to the score, is the only unit that would catch such a difference.

Is your measured cost the same as my bill?

No. We take the token counts the API reported and multiply by each model's list price on a stated date — 2026-08-06 for both Pro models, 2026-07-29 for GPT-5.6 Sol and GPT-5.4. It is a derived measured cost, not a vendor invoice. Your bill will differ with prompt length, caching, discounts, provider routing, reasoning effort and any price change since those dates. The ratios between models are the durable part; the absolute dollars are not.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.