Developer Guide

Chat Completions API: Which Parameters Actually Work (Tested on 19 Models)

The Chat Completions API is the POST /v1/chat/completions request almost every LLM provider now accepts. The documentation tells you what each parameter means. It does not tell you what happens when the model behind the endpoint does not support it. On 2026-10-03 we sent ten common parameters to 19 models through OpenRouter and checked each response. By default almost nothing was rejected, but a lot was quietly dropped. n returned one choice on all 19 models. stop was ignored on 10. And max_tokens: 5 came back as an empty string on 8, because the reasoning used the whole budget.

DataLLM Lab article cover: Chat Completions API: Which Parameters Actually Work (Tested on 19 Models)

Most guides to the Chat Completions API repeat the reference documentation. This one is the result of sending it. Every row below is a request we made on 2026-10-03, and every cell reports what came back. If you want to know what the protocol is and why other providers copied it, read our explainer on the OpenAI-compatible API first. This page is about what happens after you hit send.

The minimal request, three ways

The request needs only two fields: model and messages. Everything else is optional, and that is where the trouble starts. Here is the same call in curl and with the OpenAI Python SDK's client.chat.completions.create. Point base_url at any compatible endpoint:

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-6-luna","messages":[{"role":"user","content":"Say OK"}]}'
from openai import OpenAI
client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key=KEY)
r = client.chat.completions.create(
    model="openai/gpt-6-luna",
    messages=[{"role": "user", "content": "Say OK"}],
)
print(r.choices[0].message.content, r.choices[0].finish_reason, r.usage)

Print finish_reason and usage every time while you are developing. Both of the surprises on this page show up there first, not in the text.

The parameter matrix: 19 models × 10 parameters

Each cell has two results. The first is OpenRouter's default routing. The second is with provider: {"require_parameters": true}, which asks OpenRouter to route only to providers that support every parameter you sent. applied means we could see the effect in the response. ok means the request returned 200 and the effect cannot be seen in one call (temperature, seed, penalties). ignored means 200 but the effect did not happen. 404 means no provider met the requirement. empty means the call “worked” and returned no text.

Modeltemperaturetop_pseedstopmax_tokens 5logprobsfrequency_penalty
GPT-6 Solok / 404ok / 404ok / okignored / 40416 billed400 / 404ok / 404
GPT-6 Lunaok / 404ok / 404ok / okignored / 40416 billed400 / 404ok / 404
GPT-6.1 Solok / 404ok / 404ok / okignored / 40416 billed400 / 404ok / 404
GPT-5.4 Miniok / 404ok / 404ok / okignored / 40416 billedapplied / 404ok / 404
GPT-4.1ok / okok / okok / okignored / 40416 billedapplied / 404ok / 404
GPT-4o Miniok / okok / okok / okappliedappliedappliedok / ok
gpt-oss-120bok / okok / okok / okignored / appliedemptyappliedok / ok
Claude Sonnet 5.5ok / okok / 404ok / 404appliedappliedignored / 404ok / 404
Claude Opus 5.5ok / okok / 404ok / 404appliedappliedignored / 404ok / 404
Grok 4.7ok / okok / okok / okignored / 40444 billedignored / ignoredok / 404
Gemini 3.8 Flashok / okok / okok / okappliedemptyignored / 404ok / 404
Gemini 3.1 Pro Previewok / okok / okok / okappliedemptyignored / 404ok / 404
GLM-5.3ok / okok / okok / okignored / emptyemptyappliedok / ok
Kimi K3ok / okok / okok / okappliedemptyappliedok / ok
DeepSeek V4.1 Flashok / okok / okok / okappliedemptyappliedok / ok
Qwen3.8 Max 0902ok / okok / okok / okappliedemptyappliedok / ok
Devstral 2512ok / okok / okok / okappliedappliedignored / 404ok / ok
MiMo V2.6 Flashok / okok / okok / okignored / appliedemptyignored / appliedok / ok
Solar Mini 4ok / okok / okok / 404ignored / 404appliedignored / 404ok / ok

Two columns are missing because they were uniform. max_completion_tokens: 5 behaved exactly like max_tokens: 5 on every model. response_format: {"type": "json_object"} returned parseable JSON on 18 of 19. The exception was Claude Opus 5.5, which wrapped its JSON in a markdown code fence in both routing modes. A raw json.loads on that response fails. We tested JSON modes in depth separately in our structured outputs test.

What default routing did with each parameter, out of 19 modelsEvery bar is a count of models. All of these calls returned HTTP 200 except the three logprobs 400s.n=2 returned 1 choice19stop ignored10max_tokens 5 → empty text8logprobs ignored8temperature silently dropped4logprobs rejected (400)3Scale: width = models × 28 px. Tested 2026-10-03 via OpenRouter, default routing.
The longest bar is a parameter that never worked on any model. Every bar except the last is a 200 OK response.

Silent drops, and how to make them loud

The most useful finding is the second column of each cell. Under default routing, temperature on GPT-6 Sol returned 200. With require_parameters it returned:

404  No endpoints found that can handle the requested parameters.

So the 200 meant OpenRouter had dropped the parameter. The same thing happened for top_p on both Claude 5.5 models, for seed on Claude and Solar Mini 4, and for frequency_penalty on all five GPT-4.1 to GPT-6.1 models, Claude, Grok and Gemini. A dropped parameter does not raise an error, does not add a warning, and costs you the same.

The catalogue already says this, if you look. Each model in OpenRouter's /api/v1/models response has a supported_parameters list. In the 2026-10-03 snapshot, openai/gpt-6-sol lists no temperature, top_p or stop. All 41 of the 404s we received match a parameter that the catalogue left out for that model. The check is not complete, though: n and max_completion_tokens passed require_parameters on models whose list omits them. The useful habit is to set require_parameters: true during development. You get an error now instead of a mystery later.

One detail for anyone parsing errors: OpenRouter's error reference describes unmet routing requirements under 503. What we actually received for require_parameters was a 404. Match on the message as well as the code. Our LLM API error code guide has the rest of the codes.

max_tokens is not a hard cap

We sent max_tokens: 5 with a prompt to count to 20. Only five models did what the parameter promises: Solar Mini 4, GPT-4o Mini, both Claude 5.5 models and Devstral returned five completion tokens and stopped. The others fell into two groups:

The second group causes real outages. A classifier that asks for a one-word answer with a small max_tokens works on a non-reasoning model. Swap in a reasoning model and every response is an empty string with a success code. On reasoning models, size max_tokens for the thinking plus the answer, or lower the thinking. Our reasoning effort test measures what lowering it costs in accuracy.

The one parameter that was rejected

Under default routing, exactly one parameter produced an error. logprobs: true on GPT-6 Sol, GPT-6 Luna and GPT-6.1 Sol returned 400, wrapped by OpenRouter as “Provider returned error” around OpenAI's own message:

"message": "logprobs are not supported with reasoning models.",
"type": "invalid_request_error", "code": "unsupported_parameter"

That is the honest behaviour, and it was the exception. Claude, Gemini, Grok, Devstral, MiMo V2.6 Flash and Solar Mini 4 accepted logprobs and returned none. xAI's own models page, read 2026-10-03, confirms the Grok case: it says logprobs and top_logprobs are silently ignored on grok-4.20 and newer. If your code reads choices[0].logprobs, check that it exists before you index into it.

n and stop

n: 2 asks for two independent completions in one call. We got one choice back on all 19 models, including under require_parameters. That makes it the only parameter that was ignored in both modes everywhere. If you need several samples, send several requests.

stop: [" 7"] should cut the count off at six. It was ignored on 10 of 19 models under default routing: Solar Mini 4, all five GPT-4.1 to GPT-6.1 models, gpt-oss-120b, Grok 4.7, GLM-5.3 and MiMo V2.6 Flash. Two of those, gpt-oss-120b and MiMo, applied it once require_parameters moved the call to a provider that supports it. GLM-5.3 lists stop as supported. Under require_parameters it was served by DeepInfra and returned an empty response, so neither mode gave a usable result. Do not rely on stop to end structured output; parse it yourself.

A request checklist

  1. Read the model's supported_parameters from the catalogue before you send anything else.
  2. Develop with require_parameters: true, so that dropped parameters fail with an error.
  3. Log finish_reason and usage.completion_tokens_details.reasoning_tokens. An empty string with length means the thinking used up the budget.
  4. Do not use n. Send separate requests.
  5. Strip a markdown fence before parsing JSON, even when you set response_format.
  6. Treat temperature on GPT-5 and later as decoration. Our GPT temperature test shows it changes nothing.

If you stream, the same rules apply chunk by chunk. Our streaming guide covers where finish_reason and usage show up in a stream. Tools and function calling have their own field set, covered in our function calling guide.

How we tested

On 2026-10-03 we sent one request per model and parameter through OpenRouter, using the counting prompt Count from 1 to 20 as digits separated by single spaces. We also sent a JSON variant of the prompt for response_format. Each parameter went out twice, once with default routing and once with provider.require_parameters: true. Network or 5xx failures were retried up to three times. We recorded HTTP status, error text, the serving provider, finish_reason, token counts and the API-reported usage.cost. The whole matrix, baseline calls included, cost $0.2621. We judged the effect from the response itself. A stop counted as applied if 6 appeared and 7 did not. A max_tokens cap counted as applied if finish_reason was length or completion tokens were 5 or fewer. Calls went through OpenRouter, not through the DataLLM Lab gateway. The raw results are in our data folder as probe_params.json.

What this cannot tell you

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Primary sources, checked October 3, 2026: OpenRouter request parameters; OpenRouter provider routing. Dated measurements above may differ from the current documentation.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.