GPT Temperature on OpenRouter: Unsupported Settings and Repeatability
GPT temperature controls sampling randomness when the endpoint supports it. Our tests used OpenRouter default routing, not OpenAI's direct API. In OpenRouter's 2026-10-03 catalogue, 38 of the 40 GPT-5-and-later model IDs list no temperature parameter. Four tested IDs returned 200 with the unsupported setting under default routing, but 404 when we required parameter support. That is consistent with OpenRouter dropping unsupported settings. Separately, five identical temperature-0 calls gave the same sentence five times on 2 of 19 models; this small test is not a determinism guarantee.
Most of what is written about GPT temperature was written for GPT-3.5 and GPT-4. The definition has not changed. What has changed is whether the newest models listen. This page gives the short definition, then the measurements we made on 2026-10-03. For Anthropic's models, see our Claude temperature guide.
What temperature means
At each step a model produces a probability for every possible next token. Temperature rescales those probabilities before one is picked. At 0 the most likely token is meant to win every time. At 1 the model samples from its distribution as trained. Above 1 the unlikely tokens get a bigger share, so output becomes more varied and, past a point, less coherent. top_p is a related setting that cuts off the unlikely tail instead of reweighting it.
In plain terms: low temperature for extraction, classification and code, higher temperature for brainstorming and fiction. That advice still holds for models that accept the parameter. The rest of this page is about the ones that do not.
GPT-5 and later: accepted, then dropped
OpenRouter's model catalogue has a supported_parameters list for every model. In our 2026-10-03 snapshot, 40 OpenAI model IDs start at GPT-5 or later (not counting :batch variants). 38 of them list no temperature and no top_p. The two exceptions are openai/gpt-5-image and openai/gpt-5-image-mini. The o-series reasoning models, such as o3 and o4-mini, list no temperature either. On the older side, gpt-4.1, gpt-4o and the open-weight gpt-oss-120b still list it.
Missing from the list does not mean you get an error. We sent temperature: 1.7 to GPT-6 Sol, GPT-6 Luna, GPT-6.1 Sol and GPT-5.4 Mini under OpenRouter's default routing, and all four returned 200 OK with normal text. The parameter was dropped before it reached the model. The same happened to top_p. Our Chat Completions parameter test has the full 19-model table, including the other parameters these models drop.
| Model | Lists temperature? | temperature 1.7, default routing | Same, with require_parameters |
|---|---|---|---|
| GPT-6 Sol | No | 200, dropped | 404 |
| GPT-6 Luna | No | 200, dropped | 404 |
| GPT-6.1 Sol | No | 200, dropped | 404 |
| GPT-5.4 Mini | No | 200, dropped | 404 |
| GPT-4.1 | Yes | 200 | 200 |
| GPT-4o Mini | Yes | 200 | 200 |
| gpt-oss-120b | Yes | 200 | 200 |
How to prove it was dropped
A 200 does not tell you whether the setting was applied. Two checks do.
First, make the router refuse. Add "provider": {"require_parameters": true} to the request. OpenRouter then routes only to providers that support every parameter you sent. For all four GPT-5.4 and GPT-6 models the answer was 404 No endpoints found that can handle the requested parameters. That 404 is the evidence that the 200 was a silent drop.
Second, look for an effect, without overreading five samples. On GPT-4.1, five calls at temperature 0 gave one distinct sentence, and five at temperature 1 gave five. On GPT-6.1 Sol, temperature 0 gave five distinct sentences and temperature 1 gave three. The latter is consistent with unsupported sampling settings, but this small sample does not independently establish whether a setting was applied. The catalogue and strict-routing check are stronger evidence.
Is temperature 0 deterministic? 19 models
This is the question behind most searches for GPT temperature, so we tested it directly. We sent the prompt Write exactly one sentence describing a lighthouse at night five times to each of 19 models, with identical requests at temperature: 0. We then counted distinct outputs.
Only GPT-4.1 and Devstral 2512 produced the same sentence all five times. Gemini 3.1 Pro Preview gave 2 distinct sentences. Solar Mini 4 gave 3. GPT-6 Sol, GPT-6 Luna, GPT-4o Mini, Grok 4.7 and Qwen3.8 Max 0902 each gave 4. The other ten gave five different sentences in five calls. That ten includes both Claude 5.5 models, Gemini 3.8 Flash, DeepSeek V4.1 Flash, Kimi K3 and GLM-5.3, all of which accept temperature.
The variation was usually small. GPT-6 Sol's four sentences differed by a word or two: “casts a steady beam”, “sweeps its bright beam”, “casts its steady beam”. Claude Sonnet 5.5's five were different sentences. Small differences still matter. A test that compares output to a stored string fails on either.
Two causes are visible in our data, and others are not. Several open-weight models were served by more than one provider across the five calls. gpt-oss-120b went through four providers, and MiMo V2.6 Flash also went through four. Different hardware and kernels can produce different tokens from the same weights. But single-provider models varied too. Claude Sonnet 5.5 came from one provider every time and still gave 5 of 5. Most of these are reasoning models, and their reasoning step may not follow the temperature setting. We did not test that.
What seed does
Several APIs accept a seed so that repeated calls sample the same way. We re-ran the five calls with temperature: 0, seed: 42. It made one clear difference. Gemini 3.8 Flash went from 5 distinct outputs to 1. GPT-4.1 and Devstral 2512 stayed at 1. GPT-6 Luna and GPT-4o Mini went from 4 to 5. Most others moved by one to three in either direction. GLM-5.3 and DeepSeek V4.1 Flash went from 5 to 2, but at five samples per setting we cannot separate that from noise. Seed is worth setting where it is supported, but on most models it does not make output repeatable.
What to use instead
- For variety on the tested unsupported OpenRouter endpoints, ask for several options or a different angle in the prompt. Check endpoint support before relying on temperature.
- For less variety, constrain the output format. A JSON schema or a fixed set of labels limits what can differ. Our structured outputs test shows how far that goes, and where it gives wrong answers.
- For thinking depth, the setting that newer models do respond to is reasoning effort. Our reasoning effort test measures it on seven models.
- For repeatable tests, do not compare strings. Compare parsed fields, or run each case several times and assert on the share that passes.
- For caching, do not count on identical output either. Cache by request, not by expected response. Our prompt caching guide covers the input side.
What this means for our own benchmark
Every model on this site is scored by the same coding harness. It sends temperature: 0 with every request. On GPT-5.4, GPT-6 and other models that do not list the parameter, OpenRouter dropped that value, so those models ran at their default sampling. The harness never claimed more than a requested temperature. Even so, read our “temperature 0” as the setting we sent, not a guarantee of greedy decoding. For the models above, the scores reflect each model's normal behaviour. We have added this note to the methodology page.
How we tested
All calls ran on 2026-10-03 through OpenRouter, not through the DataLLM Lab gateway. For the parameter check, each model got one call with temperature: 1.7 under default routing and one with provider.require_parameters: true. For determinism, each of 19 models got three blocks of five identical calls: temperature 0, temperature 0 with seed: 42, and temperature 1. All used max_tokens 2000 and the same one-sentence prompt. We compared the trimmed text exactly. Every call succeeded, and the whole determinism run cost $0.2451 in API-reported usage.cost. The catalogue counts come from OpenRouter's /api/v1/models response captured the same day.
What this cannot tell you
- OpenAI's own API. We did not test it. Direct-endpoint support and error behavior cannot be inferred from a gateway result and may differ by model and reasoning setting.
- Five samples is small. 1 of 5 versus 2 of 5 is not a ranking. The clear signals are the extremes: 1 against 5.
- One prompt. A one-sentence creative prompt has many acceptable answers. Short factual prompts vary less.
- Why each model varies. We can show provider spread in some cases. We cannot see inside any provider's batching or kernels.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
Primary sources, checked October 3, 2026: OpenRouter request parameters; OpenRouter strict provider selection. Dated measurements above may differ from the current documentation.
DataLLM Lab