Hidden System Prompt Tokens: 604 In on One Model, 17,011 on Its Sibling
Two endpoints from one vendor received identical nine short prompts at an identical list price of $10 in / $50 out per 1M, and both scored 9 of 9. GPT-6 Astra reported 604 input tokens across the suite. GPT-6 Astra Pro reported 17,011. Extra reported input is often discussed as hidden system prompt tokens, but usage counts alone cannot identify the content or accounting behind it. The observed input footprint was 28 times as large, taking our list-price-derived cost from $8.19 to $35.44 per 1,000 tasks. Sakana's Fugu Ultra v2 reported 60,645, about 100 times the 604 baseline. These are recorded usage counts and derived costs, not evidence that we inspected a hidden prompt or a provider invoice.
Interpretation boundary: extra reported input is measured; its internal cause is not. References to hidden prompts below are an explanatory hypothesis, not proof that we obtained a provider system prompt.
Everyone budgets output. Input is the part you assume you control, because you wrote it. On several models in our set, you do not.
The pair: 604 in against 17,011 in
The cleanest way to see this is a pair of models that differ in as few ways as possible. GPT-6 Astra and GPT-6 Astra Pro come from the same vendor, carry the same published list price, took the same nine prompts, and returned the same score. The token counts are read back out of the usage object the API returns.
| Model | Input tokens, nine tasks | Score | List in / out per 1M | Derived cost / 1,000 tasks |
|---|---|---|---|---|
| GPT-6 Astra | 604 | 9/9 | $10 / $50 | $8.19 |
| GPT-6 Astra Pro | 17,011 | 9/9 | $10 / $50 | $35.44 |
Everything a price page publishes about these two rows is identical. Everything we measured about what they were billed for is not. Astra Pro carried 17,011 − 604 = 16,407 input tokens across the suite that we did not write, and both entries were priced on the same day, 2026-09-15, so this is not a price-drift artifact.
We cannot tell you what those tokens say. We have a count and nothing else: the usage object reports a quantity, not a payload. A fixed, provider-side preamble is the ordinary explanation for a same-vendor pair diverging like this, but we did not read a character of it and we are not going to guess. What we can say is that the quantity is real, it is billed, and it is invisible until after the response comes back.
What the rest of the field puts in front of you
One pair is an anecdote. The rest of the field is the check. Every model below received the same nine short Python task prompts and every model below scored 9 of 9. The only thing that varies is how much input the provider decided the request needed.
Put the same models next to their list prices and their derived cost, and the shape of the problem shows up.
| Model | Input tokens, nine tasks | Output tokens | List in / out per 1M | Derived cost / 1,000 tasks | Priced on |
|---|---|---|---|---|---|
| Fugu Ultra v2 | 60,645 | 7,011 | $5 / $30 | $57.06 | 2026-09-15 |
| GPT-6 Astra Pro | 17,011 | 2,977 | $10 / $50 | $35.44 | 2026-09-15 |
| Fugu Max | 4,466 | 1,791 | $2 / $6 | $2.19 | 2026-09-15 |
| Claude Fable 5.1 | 914 | 1,273 | $10 / $50 | $8.09 | 2026-09-15 |
| Ling 3.0 Flash VL | 741 | 3,299 | $0.06 / $0.18 | $0.07 | 2026-09-15 |
| DeepSeek V3.2 | 636 | 1,265 | $0.269 / $0.4 | $0.08 | 2026-07-30 |
| GPT-6 Astra | 604 | 1,353 | $10 / $50 | $8.19 | 2026-09-15 |
| GPT-5.4 mini | 604 | 969 | $0.75 / $4.5 | $0.53 | 2026-07-30 |
| Mercury 2.5 | 526 | 9,882 | $0.04 / $0.15 | $0.17 | 2026-09-15 |
Three readings worth having.
The Astra pair is not an outlier in the table. It is the same pattern the rest of the rows show, with the confounds stripped out: $8.19 against $35.44 on identical published prices, both priced 2026-09-15. Nothing on a price page distinguishes those two rows.
604 is what a short prompt actually costs. GPT-6 Astra and GPT-5.4 mini logged the same 604 input tokens across the suite — different vendors, different tiers, prices an order of magnitude apart, identical input footprint. That is the baseline. It is roughly what nine short specs weigh. Against it, Fugu Ultra v2's 60,645 − 604 = 60,041 extra tokens is not a prompt; it is a payload.
Mercury 2.5 is the opposite failure mode. 526 in, 9,882 out. Its bill is almost entirely output, which is the case people already model correctly. It still lands at $0.17 per 1,000 tasks, priced 2026-09-15. Input overhead is the axis nobody checks.
Why a price page cannot show you this
A price page publishes two numbers per model: dollars per million input tokens and dollars per million output tokens. Both are unit prices. Neither is a quantity. The quantity — how many input tokens a given request will actually be billed for — is set by the provider at serve time and is only observable after the fact, in the usage object that comes back with the response.
That is the whole gap. You can read a price page perfectly, build a cost model correctly, count your own prompt tokens to the digit, and still be wrong by an order of magnitude on input, because the provider inserted something in front of your message that you cannot see in advance and have no documented way to disable.
The overhead is a fixed block per call, so it is proportionally worst on the shortest requests — classification, routing, one-line extraction, the exact high-volume patterns where teams pick a model on list price alone. Our suite is made of short prompts, which is why it caught this at all. A benchmark of long documents would have buried a fixed prepend of this size in the noise.
The only reliable countermeasure is boring: log usage.prompt_tokens on your own traffic and compare it to what you sent. If you are running agents, where request counts are high and individual messages are short, this is the first place to look before you start trimming your own context.
When it matters, and when it does not
Be honest about the boundary. If your requests carry tens of thousands of tokens of retrieved context, a fixed provider-side prepend is a small percentage and your attention belongs elsewhere — on output volume, or on whether the long context is helping at all. If your requests are a couple of hundred tokens and you send millions of them, hidden system prompt tokens can be the largest single line in the bill, and it will not appear in any estimate built from published prices.
One thing we cannot measure cuts the other way, and it is worth stating plainly. Providers discount repeated input rather than billing it at full rate, and cached input is discounted input, not free input. A fixed block that never changes between calls is exactly the sort of thing that would sit in cache. We did not measure cache behaviour on these runs and we cannot know what discount your account receives, so our derived figures apply no cache discount at all. That makes $35.44 an upper bound on the input side rather than a prediction.
And one useful non-finding: the smallest input footprint among the models compared in this article, 512 tokens, belongs to a Granite 4.2 8B run that we exclude. Two of its nine tasks failed at the API layer after retries, so only seven were scored and the entry carries no valid score. We are not going to rank a seven-task run against a nine-task one, and neither should anyone quoting us. It is mentioned here only because within this comparison it marks roughly what nine short prompts weigh when nothing is added to them.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task — the harness reconnects only when the API itself errors, and never retries a wrong answer. Token counts are read from the usage object the API returns. Cost is derived from those measured token counts at list price on the date shown in the table, not a billing statement, and list prices move. Runs go through OpenRouter, deliberately not through our own gateway, so nothing here depends on our infrastructure and you can reproduce it without being our customer. Full method on the methodology page.
Input and output token counts in this article are suite totals across all nine tasks, not per-call figures.
What we did not measure
- What is in the injected prompt. We have a token count and nothing else. We have not read a character of it and we are not inferring its purpose.
- Whether the overhead is constant. One run, one day. A provider-side prompt can be versioned, A/B tested, or routed per-region. We did not sample it over time, which is the obvious next run.
- Whether any of it is cached. We did not capture cache-hit counts on these runs, so we cannot say how much of the prepend is billed at a discounted rate on repeat calls.
- Whether OpenRouter adds anything. Our requests reach these models through OpenRouter, so we cannot separate provider-injected prompt from anything a router might add. That is a real confound and we are naming it rather than assuming it away.
- Your cache discount. Derived cost applies none, so the input side of every figure here is an upper bound.
- Whether the padding buys anything. This is the one that should temper the whole article. Nine self-contained Python functions cannot separate a frontier model from a competent small one — every model in the table above scored 9 of 9, which mostly proves the tasks are too easy to discriminate. If Astra Pro's 16,407 extra input tokens encode scaffolding that pays off on genuinely hard, long-horizon work, our suite is structurally incapable of detecting it. What we can say is narrower and still useful: on work this size, you pay for the padding and get nothing measurable back.
- Repeat runs. One scored attempt per task. Token counts at temperature 0 are stable in our experience, but we have not proven it for these models.
FAQ
What are hidden system prompt tokens? Input tokens a provider prepends to your request server-side and bills you for. You never see them; they appear only as an inflated prompt_tokens count in the usage object the API returns.
How much overhead did you measure? Across our nine-task suite, GPT-6 Astra Pro logged 17,011 input tokens against plain GPT-6 Astra's 604 — 16,407 tokens we did not write — at the same $10 in / $50 out list price on 2026-09-15, with both models scoring 9/9.
Which model had the worst input overhead? Of the models compared in this article, Sakana Fugu Ultra v2, at 60,645 input tokens across the same nine short prompts. It is also the priciest entry in our benchmark overall, at a derived $57.06 per 1,000 tasks priced 2026-09-15.
Does caching cancel it out? It would reduce it, but we did not measure it. Providers discount repeated input rather than billing it in full, and a fixed prepend is a good candidate for that discount. We did not capture cache-hit counts on these runs, so our derived costs apply no cache discount and the input side of every figure here is an upper bound.
How do I check my own models for this? Send a request whose token count you know, then read prompt_tokens back from the response and subtract. Anything above what you sent is overhead. Do this before you finish estimating your API costs, not after.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab