GPT-6.1 Sol Review: Lower Task Cost, Not a Universal Upgrade
GPT-6.1 Sol used fewer billable output tokens on our nine Python tasks. It scored 9/9 at a derived $1.61 per 1,000 tasks; GPT-6 Sol scored 9/9 at $2.03, using the uncached $2 input / $10 output per 1M prices captured on 2026-10-02. That is about 21% lower on this run, not a promised saving on other work. Sol 6.1 also had a slower mean response: 5.9 seconds against 5.2. This review separates those dated measurements from the official API details and the checks needed before an upgrade.
The two models share the short-prompt, uncached input and output prices used in our benchmark. Their entire rate cards are not identical: cached input differs. Our short uncached requests do not test that difference. The cost gap below comes from the recorded output token counts, not a new discount.
OpenAI's API changelog dates the release of gpt-6.1-sol to September 29, 2026. The API details below were checked on October 4. The benchmark itself remains the October 2 run; updating this page did not rerun it.
The result
| Metric | GPT-6.1 Sol | GPT-6 Sol |
|---|---|---|
| Score | 9/9 | 9/9 |
| Derived cost / 1,000 tasks, priced 2026-10-02 | $1.61 | $2.03 |
| List price in / out per 1M, 2026-10-02 | $2 / $10 | $2 / $10 |
| Reasoning tokens per call | 48 | 85 |
| Input tokens across the suite | 604 | 604 |
| Output tokens across the suite | 1,329 | 1,704 |
| Mean latency | 5.9s | 5.2s |
| Rank on cost, of 75 models at 9/9 | 31 | 36 |
| Rank on latency, of 75 models at 9/9 | 31 | 26 |
| Context window | 1,050,000 | 1,050,000 |
| Measured | 2026-10-02 | 2026-10-02 |
Both models were run on the same day, through the same harness, on the same nine prompts. Neither missed a task. Among the 75 models that had scored 9 out of 9 as of 2026-10-02, GPT-6.1 Sol ranks 31st on cost and GPT-6 Sol 36th.
GPT-6.1 Sol model ID and official API pricing
OpenAI documents gpt-6.1-sol. Our benchmark used the OpenRouter ID openai/gpt-6.1-sol; do not substitute one provider's identifier into another provider's SDK. The official model page lists standard text-token rates of $2 input, $0.10 cached input, $2.50 cache write and $10 output per million for prompts up to 272K input tokens. Longer prompts, processing tiers, regional processing and tools can change the total.
The same page specifies Responses for tool calling and documents the supported reasoning efforts. These are documentation facts, not features exercised by our single-turn Python suite. Check your endpoint and parameters before changing the model: a compatible name does not establish equivalent tool behavior.
For the short, uncached requests in this suite, the calculation is ((input tokens x $2 + output tokens x $10) / 1,000,000) / 9 x 1,000. That produces a per-thousand-task estimate from nine tasks. It is not the receipt for a thousand completed requests.
Where the 42 cents went
The arithmetic is short. ($2.03 − $1.61) / $2.03 is a 21% saving per 1,000 tasks, priced 2026-10-02. Because both models consumed exactly 604 input tokens across the suite, the input side of the bill is identical to the cent. The whole gap is output.
- Output tokens: 1,704 − 1,329 = 375 fewer across the suite, which is 22% fewer. At $10 per million output tokens that is where the money goes.
- Reasoning tokens: 85 − 48 = 37 fewer per call, or 44% fewer. Reasoning tokens are billed as output, so they sit inside the 375.
Put plainly: GPT-6.1 Sol reached the same nine correct answers while thinking less before it answered and writing less when it did. On a rate card where output costs five times input, terse is cheap. We have made the same point about reasoning tokens deciding the agent bill; this is a clean case of it inside a single model family, with price held fixed.
Note what we are not claiming. A 21% gap on one run per task is a measured difference, not a guaranteed one. We do not average repeat runs, so the size could move. The direction is consistent across every token line in the table, which is why we are willing to call it.
It now spends tokens like Astra
The more interesting comparison is upward. Here is GPT-6.1 Sol beside GPT-6 Astra, which we measured in September:
| GPT-6.1 Sol | GPT-6 Astra | |
|---|---|---|
| Score | 9/9 | 9/9 |
| Input tokens across the suite | 604 | 604 |
| Output tokens across the suite | 1,329 | 1,353 |
| Reasoning tokens per call | 48 | 47 |
| Mean latency | 5.9s | 5.6s |
| List price in / out per 1M | $2 / $10 (2026-10-02) | $10 / $50 (2026-09-15) |
| Derived cost / 1,000 tasks | $1.61 (2026-10-02) | $8.19 (2026-09-15) |
On our nine prompts the two models have nearly the same token profile: 48 against 47 reasoning tokens per call, 1,329 against 1,353 output tokens. The list price differs by a factor of five ($10 / $2), and the measured cost tracks it: $8.19 / $1.61 is about 5.1x.
Those token counts are not evidence of near-Astra capability. They show how these models spent tokens on these easy tasks, not what either can do on a difficult repository. On this suite the score and verbosity were similar at very different dated prices. That makes a cheaper model worth evaluating for comparable work; it does not establish that an Astra Pro deployment is unnecessary for a different workload.
Against the rest of the field
Three comparisons matter, all priced 2026-10-02 unless stated:
- Claude Sonnet 5.5 lists at the same $2 in / $10 out and scored 9 out of 9 at $1.95. On an identical rate card, GPT-6.1 Sol came in 34 cents cheaper per 1,000 tasks. Sonnet 5.5 was faster, at 3.4s mean and 9th on latency among the 75.
- GPT-6 Sol Pro, also on the dated $2 / $10 uncached prices, had a derived cost of $7.45 for 9/9, about 4.6x GPT-6.1 Sol. Its recorded input tokens were 16,898 rather than 604 on the same prompts. That difference alone does not identify its cause. OpenRouter's public catalogue listed
openai/gpt-6.1-sol-prowhen checked on October 4; we have not measured that endpoint and do not transfer the older Pro result to it. - GPT-6 Luna scored 9 out of 9 at $0.16, roughly a tenth of GPT-6.1 Sol. Further down, Solar Mini 4 is the cheapest and fastest model to score 9 out of 9 as of 2026-10-02, at $0.03 and 2s.
That last line is the honest frame for everything above. Nine self-contained Python functions cannot separate a frontier model from a competent small one; on this suite, $0.03 and $8.19 buy the same score. What the suite can do is compare two models from the same family on the same day at the same price, which is exactly the question a point release raises. For the cheap end of the field, see the cheapest models at 9 of 9.
The one thing that got worse
GPT-6.1 Sol was slower in this run. Its mean latency was 5.9s against 5.2s for GPT-6 Sol, placing it 31st rather than 26th among the dated set of 75 models at 9/9. We did not isolate why the lower output count came with a slower mean. One run through OpenRouter on one day cannot establish a serving explanation or a stable latency penalty. GPT-6 Sol and Sonnet 5.5 were quicker on October 2; test your own route and response-time requirements.
One external signal points the same way as the cost data. In a single probe on 2026-10-02 of TypeSafe's Jev Router — which, per OpenRouter's documentation, picks a model and reasoning effort for each request — the router sent two of our nine tasks to GPT-6.1 Sol and none to GPT-6 Sol. One router, one run: a hint, not a verdict.
Should you switch?
If your GPT-6 Sol workload resembles these short coding tasks, GPT-6.1 Sol is a reasonable candidate for a controlled comparison. The uncached prices matched and the dated derived cost was lower, but a production migration is not proven by nine passing functions.
- Keep a baseline. Save representative inputs, expected outputs, current route, parameters, token usage and actual charges. Include difficult requests and failure cases, not only examples that already pass.
- Check compatibility first. Confirm the provider-specific ID, endpoint, tools, structured output and reasoning settings your application uses. Do not silently change several settings during the model comparison.
- Repeat both models. Compare accepted outputs, retries, end-to-end latency and total spend under the same conditions. Count timeouts and incorrect answers, not just successful calls.
- Decide against your requirements. Keep the new model only if it meets the quality and response-time limits you set before testing. Use a limited rollout with a working fallback rather than switching all traffic on this article's score.
If you are choosing fresh for work this size, the decision is not between Sol versions. It is whether you need a $2 / $10 model at all when GPT-6 Luna scored the same at $0.16. Pick GPT-6.1 Sol when your own workload shows that its headroom is worth ten times the bill. Earlier Sol generations are covered in our GPT-5.6 Sol review.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; neither Sol model had one. Cost is derived from measured input and output token counts multiplied by the list price captured 2026-10-02 — it is not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. The Jev Router figure comes from a separate probe on the same nine tasks. Full method on the methodology page.
What we did not measure
- Agentic coding. Nine single-turn functions cannot test multi-step repository work or tool use.
- Whether 48 reasoning tokens holds on hard problems. A model that thinks less on easy tasks may think just as much on difficult ones, and then the 21% evaporates.
- Reasoning effort settings. We did not vary them. A different effort setting could narrow or reverse the token gap.
- Cached input and cache writes. The official rates differ from uncached input, but our short prompts do not measure cache savings or cache-write charges.
- The 1,050,000-token context. Our prompts are short.
- Repeat runs. One scored attempt per task. The 21% and the 0.7 seconds are single-run readings, not averages.
- GPT-6.1 Sol Pro. Its OpenRouter catalogue listing is not a benchmark result. The older GPT-6 Sol Pro cost 3.7x GPT-6 Sol on our suite; that is not a prediction for the new endpoint.
GPT-6.1 Sol FAQ
Which model ID should I use?
OpenAI documents gpt-6.1-sol; our OpenRouter benchmark used openai/gpt-6.1-sol. Use the ID and endpoint documented by the provider you actually call. The plain and Pro catalogue entries are separate.
Does the $1.61 figure include every production charge?
No. It is a derived per-thousand-task estimate using measured tokens from nine short tasks and dated uncached prices. It does not measure a thousand requests, cache effects, long-context tiers, tools, retries or your gateway's bill.
Is GPT-6.1 Sol better than Astra or Sonnet 5.5?
This suite cannot answer that generally. These models all passed the small task set, while token usage, dated cost and mean latency differed. Test the hard work your application must complete before choosing between them.
A candidate for a controlled comparison
The useful finding in this GPT-6.1 Sol review is narrow: it passed the same nine functions as GPT-6 Sol with fewer output tokens and a lower derived task cost on October 2. Use that result to design a comparison on your own work, not as a blanket upgrade instruction.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
API sources checked October 4, 2026: OpenAI changelog; OpenAI model and pricing documentation; OpenRouter public model catalogue. Catalogue visibility is not evidence that this page tested a model or that every account has access.
DataLLM Lab