Chat Completions API: Which Parameters Actually Work (Tested on 19 Models)
The Chat Completions API is the POST /v1/chat/completions request almost every LLM provider now accepts. The documentation tells you what each parameter means. It does not tell you what happens when the model behind the endpoint does not support it. On 2026-10-03 we sent ten common parameters to 19 models through OpenRouter and checked each response. By default almost nothing was rejected, but a lot was quietly dropped. n returned one choice on all 19 models. stop was ignored on 10. And max_tokens: 5 came back as an empty string on 8, because the reasoning used the whole budget.
Most guides to the Chat Completions API repeat the reference documentation. This one is the result of sending it. Every row below is a request we made on 2026-10-03, and every cell reports what came back. If you want to know what the protocol is and why other providers copied it, read our explainer on the OpenAI-compatible API first. This page is about what happens after you hit send.
The minimal request, three ways
The request needs only two fields: model and messages. Everything else is optional, and that is where the trouble starts. Here is the same call in curl and with the OpenAI Python SDK's client.chat.completions.create. Point base_url at any compatible endpoint:
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"model":"openai/gpt-6-luna","messages":[{"role":"user","content":"Say OK"}]}'
from openai import OpenAI
client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key=KEY)
r = client.chat.completions.create(
model="openai/gpt-6-luna",
messages=[{"role": "user", "content": "Say OK"}],
)
print(r.choices[0].message.content, r.choices[0].finish_reason, r.usage)
Print finish_reason and usage every time while you are developing. Both of the surprises on this page show up there first, not in the text.
The parameter matrix: 19 models × 10 parameters
Each cell has two results. The first is OpenRouter's default routing. The second is with provider: {"require_parameters": true}, which asks OpenRouter to route only to providers that support every parameter you sent. applied means we could see the effect in the response. ok means the request returned 200 and the effect cannot be seen in one call (temperature, seed, penalties). ignored means 200 but the effect did not happen. 404 means no provider met the requirement. empty means the call “worked” and returned no text.
| Model | temperature | top_p | seed | stop | max_tokens 5 | logprobs | frequency_penalty |
|---|---|---|---|---|---|---|---|
| GPT-6 Sol | ok / 404 | ok / 404 | ok / ok | ignored / 404 | 16 billed | 400 / 404 | ok / 404 |
| GPT-6 Luna | ok / 404 | ok / 404 | ok / ok | ignored / 404 | 16 billed | 400 / 404 | ok / 404 |
| GPT-6.1 Sol | ok / 404 | ok / 404 | ok / ok | ignored / 404 | 16 billed | 400 / 404 | ok / 404 |
| GPT-5.4 Mini | ok / 404 | ok / 404 | ok / ok | ignored / 404 | 16 billed | applied / 404 | ok / 404 |
| GPT-4.1 | ok / ok | ok / ok | ok / ok | ignored / 404 | 16 billed | applied / 404 | ok / 404 |
| GPT-4o Mini | ok / ok | ok / ok | ok / ok | applied | applied | applied | ok / ok |
| gpt-oss-120b | ok / ok | ok / ok | ok / ok | ignored / applied | empty | applied | ok / ok |
| Claude Sonnet 5.5 | ok / ok | ok / 404 | ok / 404 | applied | applied | ignored / 404 | ok / 404 |
| Claude Opus 5.5 | ok / ok | ok / 404 | ok / 404 | applied | applied | ignored / 404 | ok / 404 |
| Grok 4.7 | ok / ok | ok / ok | ok / ok | ignored / 404 | 44 billed | ignored / ignored | ok / 404 |
| Gemini 3.8 Flash | ok / ok | ok / ok | ok / ok | applied | empty | ignored / 404 | ok / 404 |
| Gemini 3.1 Pro Preview | ok / ok | ok / ok | ok / ok | applied | empty | ignored / 404 | ok / 404 |
| GLM-5.3 | ok / ok | ok / ok | ok / ok | ignored / empty | empty | applied | ok / ok |
| Kimi K3 | ok / ok | ok / ok | ok / ok | applied | empty | applied | ok / ok |
| DeepSeek V4.1 Flash | ok / ok | ok / ok | ok / ok | applied | empty | applied | ok / ok |
| Qwen3.8 Max 0902 | ok / ok | ok / ok | ok / ok | applied | empty | applied | ok / ok |
| Devstral 2512 | ok / ok | ok / ok | ok / ok | applied | applied | ignored / 404 | ok / ok |
| MiMo V2.6 Flash | ok / ok | ok / ok | ok / ok | ignored / applied | empty | ignored / applied | ok / ok |
| Solar Mini 4 | ok / ok | ok / ok | ok / 404 | ignored / 404 | applied | ignored / 404 | ok / ok |
Two columns are missing because they were uniform. max_completion_tokens: 5 behaved exactly like max_tokens: 5 on every model. response_format: {"type": "json_object"} returned parseable JSON on 18 of 19. The exception was Claude Opus 5.5, which wrapped its JSON in a markdown code fence in both routing modes. A raw json.loads on that response fails. We tested JSON modes in depth separately in our structured outputs test.
Silent drops, and how to make them loud
The most useful finding is the second column of each cell. Under default routing, temperature on GPT-6 Sol returned 200. With require_parameters it returned:
404 No endpoints found that can handle the requested parameters.
So the 200 meant OpenRouter had dropped the parameter. The same thing happened for top_p on both Claude 5.5 models, for seed on Claude and Solar Mini 4, and for frequency_penalty on all five GPT-4.1 to GPT-6.1 models, Claude, Grok and Gemini. A dropped parameter does not raise an error, does not add a warning, and costs you the same.
The catalogue already says this, if you look. Each model in OpenRouter's /api/v1/models response has a supported_parameters list. In the 2026-10-03 snapshot, openai/gpt-6-sol lists no temperature, top_p or stop. All 41 of the 404s we received match a parameter that the catalogue left out for that model. The check is not complete, though: n and max_completion_tokens passed require_parameters on models whose list omits them. The useful habit is to set require_parameters: true during development. You get an error now instead of a mystery later.
One detail for anyone parsing errors: OpenRouter's error reference describes unmet routing requirements under 503. What we actually received for require_parameters was a 404. Match on the message as well as the code. Our LLM API error code guide has the rest of the codes.
max_tokens is not a hard cap
We sent max_tokens: 5 with a prompt to count to 20. Only five models did what the parameter promises: Solar Mini 4, GPT-4o Mini, both Claude 5.5 models and Devstral returned five completion tokens and stopped. The others fell into two groups:
- Billed more than you asked for. The five OpenAI GPT-4.1 to GPT-6.1 models each reported 16 completion tokens with
finish_reason: length. That looks like a floor. Grok 4.7 reported 44 completion tokens, 39 of them reasoning, and returned three numbers. - Returned nothing. Eight reasoning models spent the budget on reasoning and returned an empty
content: gpt-oss-120b, Gemini 3.8 Flash, Gemini 3.1 Pro Preview, GLM-5.3, Kimi K3, DeepSeek V4.1 Flash, Qwen3.8 Max 0902 and MiMo V2.6 Flash. The status was 200 andfinish_reasonwaslength.
The second group causes real outages. A classifier that asks for a one-word answer with a small max_tokens works on a non-reasoning model. Swap in a reasoning model and every response is an empty string with a success code. On reasoning models, size max_tokens for the thinking plus the answer, or lower the thinking. Our reasoning effort test measures what lowering it costs in accuracy.
The one parameter that was rejected
Under default routing, exactly one parameter produced an error. logprobs: true on GPT-6 Sol, GPT-6 Luna and GPT-6.1 Sol returned 400, wrapped by OpenRouter as “Provider returned error” around OpenAI's own message:
"message": "logprobs are not supported with reasoning models.",
"type": "invalid_request_error", "code": "unsupported_parameter"
That is the honest behaviour, and it was the exception. Claude, Gemini, Grok, Devstral, MiMo V2.6 Flash and Solar Mini 4 accepted logprobs and returned none. xAI's own models page, read 2026-10-03, confirms the Grok case: it says logprobs and top_logprobs are silently ignored on grok-4.20 and newer. If your code reads choices[0].logprobs, check that it exists before you index into it.
n and stop
n: 2 asks for two independent completions in one call. We got one choice back on all 19 models, including under require_parameters. That makes it the only parameter that was ignored in both modes everywhere. If you need several samples, send several requests.
stop: [" 7"] should cut the count off at six. It was ignored on 10 of 19 models under default routing: Solar Mini 4, all five GPT-4.1 to GPT-6.1 models, gpt-oss-120b, Grok 4.7, GLM-5.3 and MiMo V2.6 Flash. Two of those, gpt-oss-120b and MiMo, applied it once require_parameters moved the call to a provider that supports it. GLM-5.3 lists stop as supported. Under require_parameters it was served by DeepInfra and returned an empty response, so neither mode gave a usable result. Do not rely on stop to end structured output; parse it yourself.
A request checklist
- Read the model's
supported_parametersfrom the catalogue before you send anything else. - Develop with
require_parameters: true, so that dropped parameters fail with an error. - Log
finish_reasonandusage.completion_tokens_details.reasoning_tokens. An empty string withlengthmeans the thinking used up the budget. - Do not use
n. Send separate requests. - Strip a markdown fence before parsing JSON, even when you set
response_format. - Treat
temperatureon GPT-5 and later as decoration. Our GPT temperature test shows it changes nothing.
If you stream, the same rules apply chunk by chunk. Our streaming guide covers where finish_reason and usage show up in a stream. Tools and function calling have their own field set, covered in our function calling guide.
How we tested
On 2026-10-03 we sent one request per model and parameter through OpenRouter, using the counting prompt Count from 1 to 20 as digits separated by single spaces. We also sent a JSON variant of the prompt for response_format. Each parameter went out twice, once with default routing and once with provider.require_parameters: true. Network or 5xx failures were retried up to three times. We recorded HTTP status, error text, the serving provider, finish_reason, token counts and the API-reported usage.cost. The whole matrix, baseline calls included, cost $0.2621. We judged the effect from the response itself. A stop counted as applied if 6 appeared and 7 did not. A max_tokens cap counted as applied if finish_reason was length or completion tokens were 5 or fewer. Calls went through OpenRouter, not through the DataLLM Lab gateway. The raw results are in our data folder as probe_params.json.
What this cannot tell you
- Direct provider APIs. We called OpenAI-compatible routes on OpenRouter. OpenAI's, Google's or Anthropic's native APIs may reject what OpenRouter drops.
- Effects you cannot see in one call. “ok” for temperature, seed and penalties only means the call succeeded. Whether temperature changed anything is a separate measurement, which our temperature test makes.
- Provider churn. Which provider serves a model changes over time. That also changes which parameters survive. Re-check before you depend on one.
- One request per cell. A provider that fails intermittently could have landed on either side of our result.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
Primary sources, checked October 3, 2026: OpenRouter request parameters; OpenRouter provider routing. Dated measurements above may differ from the current documentation.
DataLLM Lab