GPT-6 Pro: Extra Reported Input, the Same Nine Passing Tasks
Every GPT-6 Pro endpoint in this run reported roughly 28 times the input tokens of its standard sibling on identical nine prompts. GPT-6 Sol, Luna and Astra each logged 604 input tokens across our suite. Sol Pro logged 16,898, Luna Pro 17,806, Astra Pro 17,011, whether the model lists at $0.1 or $10 per million input tokens. All six scored 9 out of 9. Reasoning tokens per call rose by at most 2.36x (Astra Pro, 111 ÷ 47). In our derived cost calculation, extra reported input explains most of the short-task Pro premium. We cannot identify a hidden prompt or infer a fixed overhead for every request from these usage counts.
Interpretation boundary: prompt overhead below is a possible explanation for extra reported input, not an inspected system prompt. Provider wrappers or usage accounting may also contribute. The measurements do not establish the internal cause.
When we measured Astra Pro in September, one data point could have been a quirk of one endpoint. With Sol Pro and Luna Pro measured on 2026-10-02, it is a pattern across the whole tier.
Three Pro endpoints, one input signature
The usual assumption is that a Pro endpoint costs more because it thinks longer. Here is what the same nine prompts actually logged:
| Endpoint | Score | Measured cost / 1k tasks | Priced at (in / out per 1M) | Input tokens | Output tokens | Reasoning / call | Mean latency |
|---|---|---|---|---|---|---|---|
| GPT-6 Sol | 9/9 | $2.03 | $2 / $10, 2026-10-02 | 604 | 1,704 | 85 | 5.2s |
| GPT-6 Sol Pro | 9/9 | $7.45 | $2 / $10, 2026-10-02 | 16,898 | 3,322 | 158 | 5.9s |
| GPT-6 Luna | 9/9 | $0.16 | $0.1 / $0.5, 2026-10-02 | 604 | 2,763 | 201 | 4.9s |
| GPT-6 Luna Pro | 9/9 | $0.49 | $0.1 / $0.5, 2026-10-02 | 17,806 | 5,248 | 370 | 7.3s |
| GPT-6 Astra | 9/9 | $8.19 | $10 / $50, 2026-09-15 | 604 | 1,353 | 47 | 5.6s |
| GPT-6 Astra Pro | 9/9 | $35.44 | $10 / $50, 2026-09-15 | 17,011 | 2,977 | 111 | 8.4s |
Sol and Luna were run on 2026-10-02 and priced the same day. Astra and Astra Pro were run on 2026-09-16 at prices captured 2026-09-15. Each Pro tier lists at exactly the same per-token price as its base — the rate card shows no premium at all. The premium appears only once you count tokens.
Three things stand out. First, the three base models agree to the token: 604 input each, because they read the identical prompts and add nothing. Second, the three Pro endpoints land within 908 tokens of each other — 17,806 − 16,898 — across nine requests, on models whose input prices span $0.1 to $10 per million. A block that size, that stable, across that price range, behaves like shared scaffolding rather than model-specific deliberation. Third, input is the column that moved most: 28x for Sol Pro (16,898 ÷ 604), 29x for Luna Pro (17,806 ÷ 604), 28x for Astra Pro (17,011 ÷ 604).
One request, read raw
Suite totals hide where the tokens sit, so in September we sent a single 65-token message to Astra and to Astra Pro and read the usage payload back. That probe is published in our GPT-6 Astra review:
| Endpoint | Prompt tokens billed | of which cached |
|---|---|---|
| GPT-6 Astra | 65 | — |
| GPT-6 Astra Pro | 1,721 | 1,389 |
Same 65-token message. Astra billed exactly what we sent. Astra Pro billed 1,656 more (1,721 − 65) — prompt tokens added before ours, on a request where no reasoning had happened yet. That is the mechanism the suite totals point at: a block of input attached to every call, independent of what you ask. We wrote up the general phenomenon, and how far other vendors take it, in our piece on hidden system prompt tokens.
We have not re-run the single-request probe on Sol Pro or Luna Pro. What we can say is that their suite totals sit in the same narrow band as Astra Pro's, which is what a shared fixed block would produce and what model-specific extra reasoning would not.
Prompt or thinking: where the premium goes
The honest answer is both, in unequal shares. Here is the token difference per pair, every figure a plain subtraction from the table above:
| Pair | Cost premium / 1k tasks | Extra input tokens | Extra output tokens | Cost ratio |
|---|---|---|---|---|
| Sol → Sol Pro (priced 2026-10-02) | $5.42 | 16,294 | 1,618 | 3.7x |
| Luna → Luna Pro (priced 2026-10-02) | $0.33 | 17,202 | 2,485 | 3.1x |
| Astra → Astra Pro (priced 2026-09-15) | $27.25 | 16,407 | 1,624 | 4.3x |
In every pair the extra input outnumbers the extra output by thousands of tokens. Output is priced at five times input on all three rate cards ($10 ÷ $2, $0.5 ÷ $0.1, $50 ÷ $10), so each extra output token weighs five input tokens. Even with that weighting, the extra input is the larger term in all three pairs. It is closest to even on Luna Pro, which also wrote the most extra output — 5,248 output tokens against Luna's 2,763 — and nearly doubled its reasoning, 201 to 370 per call.
So the usual assumption is not wrong; Pro does reason more. It is incomplete for short prompts. On our tasks, the bigger cost driver is input the endpoint adds, and a price page that shows identical per-token rates for both tiers cannot tell you that.
The same signature appears outside OpenAI's Pro tier, smaller: Grok 4.7 logged 11,761 input tokens on these prompts against 2,437 for Grok 4.6. That is a version change rather than a mode switch, so we do not count it as part of this pattern.
What this means for choosing a tier
The overhead is fixed in tokens, so its weight depends on your prompt length. On our prompts — 604 input tokens across nine — the added block dwarfs the real input. On a request that already carries tens of thousands of tokens of code or documents, the same block would be a rounding error and the reasoning increase would be most of the premium. We did not test that case; it follows from the arithmetic, not from a measurement.
The tier choice outweighs the Pro switch. Luna Pro at $0.49 per 1,000 tasks (priced 2026-10-02) cost less than plain Sol at $2.03, and Sol Pro at $7.45 cost less than plain Astra at $8.19 (priced 2026-09-15). All four scored 9 out of 9. As of 2026-10-02, Luna Pro ranks 16th cheapest of the 75 models that clear our suite; Astra Pro ranks 74th.
Nothing in our suite justified any Pro tier. All six endpoints scored 9 out of 9. That is a statement about our suite, not about Pro: nine self-contained Python functions cannot separate a frontier model from a competent small one. Solar Mini 4 cleared the same nine at $0.03 per 1,000 tasks (priced 2026-10-02) in a 2-second mean, the cheapest and fastest model to score 9 out of 9 in our set as of 2026-10-02. If Pro buys accuracy, it buys it on problems harder than these, and you should measure that on your own workload before paying 3.1x to 4.3x for it. For cheaper routes to the same score, see our cheap coding roundup.
Caching softens the overhead, if your traffic allows it. 1,389 of Astra Pro's 1,721 probe tokens were cached. Repeated calls with a stable prefix may pay a discounted rate on most of the added block; our prompt caching guide covers how that works. Our cost figures do not apply any discount.
How these numbers were produced
Nine Python tasks, each given as a function signature plus a spec and no example tests. The generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, so a provider error never counts as a model mistake. Cost is derived from the measured input and output token counts at the list price captured on the date shown beside each figure; it is not a billing statement, and prices move. All runs go through OpenRouter, not through the DataLLM Lab gateway. The single-request probe read the raw usage payload of one call per endpoint. Full method on the methodology page.
What we did not measure
- The probe on Sol Pro and Luna Pro. The 65-against-1,721 reading is Astra only. For the other two we infer a fixed block from suite totals, which is a strong hint, not a direct reading.
- What the added prompt contains. We measured its size, never its text.
- Long prompts. Every claim about the overhead shrinking on large inputs is arithmetic. We ran nothing longer than our short task specs.
- Cache discounts. We priced every input token at full list. A cache-heavy production bill for a Pro endpoint could sit well below our derived figure.
- Whether pro mode buys accuracy. Six 9-out-of-9 scores cannot show it. Our suite has no headroom for that question.
- Multi-turn sessions. If the block is added per request, an agent loop pays it on every turn. We tested single-turn calls only.
- GPT-6.1 Pro variants and batch endpoints. We measured GPT-6.1 Sol but no Pro variant of it, and no batch endpoint.
- Repeat runs. One scored attempt per task. Single-run figures, not averages.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab