Benchmarks

Hidden System Prompt Tokens: 604 In on One Model, 17,011 on Its Sibling

Two endpoints from one vendor received identical nine short prompts at an identical list price of $10 in / $50 out per 1M, and both scored 9 of 9. GPT-6 Astra reported 604 input tokens across the suite. GPT-6 Astra Pro reported 17,011. Extra reported input is often discussed as hidden system prompt tokens, but usage counts alone cannot identify the content or accounting behind it. The observed input footprint was 28 times as large, taking our list-price-derived cost from $8.19 to $35.44 per 1,000 tasks. Sakana's Fugu Ultra v2 reported 60,645, about 100 times the 604 baseline. These are recorded usage counts and derived costs, not evidence that we inspected a hidden prompt or a provider invoice.

Interpretation boundary: extra reported input is measured; its internal cause is not. References to hidden prompts below are an explanatory hypothesis, not proof that we obtained a provider system prompt.

DataLLM Lab article cover: Hidden System Prompt Tokens: 604 In on One Model, 17,011 on Its Sibling

Everyone budgets output. Input is the part you assume you control, because you wrote it. On several models in our set, you do not.

The pair: 604 in against 17,011 in

The cleanest way to see this is a pair of models that differ in as few ways as possible. GPT-6 Astra and GPT-6 Astra Pro come from the same vendor, carry the same published list price, took the same nine prompts, and returned the same score. The token counts are read back out of the usage object the API returns.

ModelInput tokens, nine tasksScoreList in / out per 1MDerived cost / 1,000 tasks
GPT-6 Astra6049/9$10 / $50$8.19
GPT-6 Astra Pro17,0119/9$10 / $50$35.44

Everything a price page publishes about these two rows is identical. Everything we measured about what they were billed for is not. Astra Pro carried 17,011 − 604 = 16,407 input tokens across the suite that we did not write, and both entries were priced on the same day, 2026-09-15, so this is not a price-drift artifact.

We cannot tell you what those tokens say. We have a count and nothing else: the usage object reports a quantity, not a payload. A fixed, provider-side preamble is the ordinary explanation for a same-vendor pair diverging like this, but we did not read a character of it and we are not going to guess. What we can say is that the quantity is real, it is billed, and it is invisible until after the response comes back.

What the rest of the field puts in front of you

One pair is an anecdote. The rest of the field is the check. Every model below received the same nine short Python task prompts and every model below scored 9 of 9. The only thing that varies is how much input the provider decided the request needed.

Same nine prompts. Same 9/9. Wildly different input bills.Input tokens billed across the whole suite. The prompts were identical; the padding in front of them was not.Fugu Ultra v260,645GPT-6 Astra Pro17,011Fugu Max4,466Claude Fable 5.1914Ling 3.0 Flash VL741GPT-6 Astra604Mercury 2.5526One scale throughout: 8.5 px per 1,000 input tokens. All seven models scored 9/9 on the identical nine executed Python tasks.
The bottom four bars are not a rendering bug. At one honest scale, that is what a normal input footprint looks like next to Fugu Ultra v2.

Put the same models next to their list prices and their derived cost, and the shape of the problem shows up.

ModelInput tokens, nine tasksOutput tokensList in / out per 1MDerived cost / 1,000 tasksPriced on
Fugu Ultra v260,6457,011$5 / $30$57.062026-09-15
GPT-6 Astra Pro17,0112,977$10 / $50$35.442026-09-15
Fugu Max4,4661,791$2 / $6$2.192026-09-15
Claude Fable 5.19141,273$10 / $50$8.092026-09-15
Ling 3.0 Flash VL7413,299$0.06 / $0.18$0.072026-09-15
DeepSeek V3.26361,265$0.269 / $0.4$0.082026-07-30
GPT-6 Astra6041,353$10 / $50$8.192026-09-15
GPT-5.4 mini604969$0.75 / $4.5$0.532026-07-30
Mercury 2.55269,882$0.04 / $0.15$0.172026-09-15

Three readings worth having.

The Astra pair is not an outlier in the table. It is the same pattern the rest of the rows show, with the confounds stripped out: $8.19 against $35.44 on identical published prices, both priced 2026-09-15. Nothing on a price page distinguishes those two rows.

604 is what a short prompt actually costs. GPT-6 Astra and GPT-5.4 mini logged the same 604 input tokens across the suite — different vendors, different tiers, prices an order of magnitude apart, identical input footprint. That is the baseline. It is roughly what nine short specs weigh. Against it, Fugu Ultra v2's 60,645 − 604 = 60,041 extra tokens is not a prompt; it is a payload.

Mercury 2.5 is the opposite failure mode. 526 in, 9,882 out. Its bill is almost entirely output, which is the case people already model correctly. It still lands at $0.17 per 1,000 tasks, priced 2026-09-15. Input overhead is the axis nobody checks.

Why a price page cannot show you this

A price page publishes two numbers per model: dollars per million input tokens and dollars per million output tokens. Both are unit prices. Neither is a quantity. The quantity — how many input tokens a given request will actually be billed for — is set by the provider at serve time and is only observable after the fact, in the usage object that comes back with the response.

That is the whole gap. You can read a price page perfectly, build a cost model correctly, count your own prompt tokens to the digit, and still be wrong by an order of magnitude on input, because the provider inserted something in front of your message that you cannot see in advance and have no documented way to disable.

The overhead is a fixed block per call, so it is proportionally worst on the shortest requests — classification, routing, one-line extraction, the exact high-volume patterns where teams pick a model on list price alone. Our suite is made of short prompts, which is why it caught this at all. A benchmark of long documents would have buried a fixed prepend of this size in the noise.

The only reliable countermeasure is boring: log usage.prompt_tokens on your own traffic and compare it to what you sent. If you are running agents, where request counts are high and individual messages are short, this is the first place to look before you start trimming your own context.

When it matters, and when it does not

Be honest about the boundary. If your requests carry tens of thousands of tokens of retrieved context, a fixed provider-side prepend is a small percentage and your attention belongs elsewhere — on output volume, or on whether the long context is helping at all. If your requests are a couple of hundred tokens and you send millions of them, hidden system prompt tokens can be the largest single line in the bill, and it will not appear in any estimate built from published prices.

One thing we cannot measure cuts the other way, and it is worth stating plainly. Providers discount repeated input rather than billing it at full rate, and cached input is discounted input, not free input. A fixed block that never changes between calls is exactly the sort of thing that would sit in cache. We did not measure cache behaviour on these runs and we cannot know what discount your account receives, so our derived figures apply no cache discount at all. That makes $35.44 an upper bound on the input side rather than a prediction.

And one useful non-finding: the smallest input footprint among the models compared in this article, 512 tokens, belongs to a Granite 4.2 8B run that we exclude. Two of its nine tasks failed at the API layer after retries, so only seven were scored and the entry carries no valid score. We are not going to rank a seven-task run against a nine-task one, and neither should anyone quoting us. It is mentioned here only because within this comparison it marks roughly what nine short prompts weigh when nothing is added to them.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task — the harness reconnects only when the API itself errors, and never retries a wrong answer. Token counts are read from the usage object the API returns. Cost is derived from those measured token counts at list price on the date shown in the table, not a billing statement, and list prices move. Runs go through OpenRouter, deliberately not through our own gateway, so nothing here depends on our infrastructure and you can reproduce it without being our customer. Full method on the methodology page.

Input and output token counts in this article are suite totals across all nine tasks, not per-call figures.

What we did not measure

FAQ

What are hidden system prompt tokens? Input tokens a provider prepends to your request server-side and bills you for. You never see them; they appear only as an inflated prompt_tokens count in the usage object the API returns.

How much overhead did you measure? Across our nine-task suite, GPT-6 Astra Pro logged 17,011 input tokens against plain GPT-6 Astra's 604 — 16,407 tokens we did not write — at the same $10 in / $50 out list price on 2026-09-15, with both models scoring 9/9.

Which model had the worst input overhead? Of the models compared in this article, Sakana Fugu Ultra v2, at 60,645 input tokens across the same nine short prompts. It is also the priciest entry in our benchmark overall, at a derived $57.06 per 1,000 tasks priced 2026-09-15.

Does caching cancel it out? It would reduce it, but we did not measure it. Providers discount repeated input rather than billing it in full, and a fixed prepend is a good candidate for that discount. We did not capture cache-hit counts on these runs, so our derived costs apply no cache discount and the input side of every figure here is an upper bound.

How do I check my own models for this? Send a request whose token count you know, then read prompt_tokens back from the response and subtract. Anything above what you sent is overhead. Do this before you finish estimating your API costs, not after.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.