Developer Guide

OpenAI Structured Outputs: Valid JSON Is Not the Same as a Right Answer (Tested)

Test scope: synthetic invoice fixtures through OpenRouter, not a direct-vendor endpoint test. OpenAI structured outputs promise JSON that matches your schema, and on our test they delivered. Under strict: true all six OpenAI models we tried returned schema-valid JSON on all 8 of our deliberately awkward invoices. What the schema cannot promise is that the values are true. When an invoice was in Swiss francs and the schema only allowed USD, EUR, GBP or JPY, all six models picked one of the four. The JSON was valid, and the currency was wrong. When we redesigned the schema so that the true answer had somewhere to go, five of the six scored 8 out of 8.

DataLLM Lab article cover: OpenAI Structured Outputs: Valid JSON Is Not the Same as a Right Answer (Tested)

Structured outputs solve a real problem. Before them, a model could return a broken bracket, a trailing comment or a field you never asked for, and your parser would crash at 3am. Strict mode removes that whole class of failure. This page is about the failure it does not remove. We ran the same invoice extraction three ways on six OpenAI models on 2026-10-03, scored the answers as well as the shape, and then tested a fix. If you use Claude, our Claude structured output guide covers the Anthropic side. The Gemini results are in a separate Gemini test.

Three ways to ask for JSON

Our schema asked for vendor, date (YYYY-MM-DD), currency (an enum of USD, EUR, GBP, JPY), total, paid (boolean), and line_items, each with description, qty (integer) and unit_price. additionalProperties was false everywhere. It is a typical first-draft schema. That is the point.

Clean invoices: everyone passes

We first ran 10 ordinary invoices: different languages and date formats, USD, EUR, GBP and JPY, paid and unpaid. All six models returned schema-valid JSON with the right total, currency, date, paid flag and item count on all 10, in all three modes. One thing differed. In prompt-only mode, GPT-4.1 Mini wrapped its JSON in a markdown code fence on all 10. A plain json.loads fails on every one of those. The other five returned bare JSON. On clean input, strict mode's only advantage is that you never have to strip a fence.

Eight awkward invoices

Then we wrote eight invoices that break the first-draft schema on purpose. Each one tests a single thing:

CaseThe catchRight answer
h01Hotel bill in CHF, which the enum does not allowCHF, or admit it does not fit
h02No date anywhere on the documentEmpty or null, not an invented date
h032.5 and 1.5 consulting hours, but qty must be an integerLine items that still add up to the $725 total
h0440 line items40 items, total $1,594
h05Quotes, backslashes and a tab inside descriptions3 items, valid JSON
h0603/04/2026 from a London vendor2026-04-03
h07A $20 credit note, so the amount due is not the item sumTotal $100
h0850% deposit paid, balance outstandingpaid false, total $2,400
ModelStrict: schema-validStrict: rightjson_object: rightPrompt only: right
GPT-6 Sol8/87/88/88/8
GPT-6 Luna8/87/88/88/8
GPT-6.1 Sol8/87/88/88/8
GPT-5.4 Mini8/86/86/86/8
GPT-4.1 Mini8/86/86/87/8
gpt-oss-120b8/86/86/88/8

Read it across. Strict mode had the best schema-validity of the three modes and, for the GPT-6 models, the worst correctness. The looser modes were right more often because the model could write "currency": "CHF" or "qty": 2.5. That broke the schema and kept the truth. A validator catches a schema break. Nothing catches a valid lie.

What strict mode did with each one

h01, the franc. Under strict mode GPT-6 Sol, GPT-5.4 Mini, GPT-4.1 Mini and gpt-oss-120b wrote USD. GPT-6 Luna and GPT-6.1 Sol wrote EUR. The total stayed at 610 on all six, so the amount was right and the currency was wrong. Without the schema, most models wrote CHF or null. GPT-5.4 Mini was the exception: it picked EUR under json_object and USD with the prompt alone, even though nothing forced it to.

h02, the missing date. All six returned an empty string under strict mode. That is the honest answer the schema allowed, since it required a string. Without the schema, all six returned null. No OpenAI model made up a date. Three Gemini models did, as the Gemini test shows.

h03, the fractional hours. This one splits the field. GPT-6 Sol, Luna and 6.1 Sol set each qty to 1 and moved the money into unit_price, $500 and $225. The line items still sum to $725. The others rounded the hours and kept the hourly rate. GPT-5.4 Mini wrote 3 and 2 hours, so its items sum to $900. GPT-4.1 Mini wrote 2 and 2, which comes to $700. gpt-oss-120b wrote 2 and 1, which comes to $550. All of those are schema-valid. Each reports a total of $725 next to line items that do not add up to it.

h04 to h08. Under strict mode all six got 40 items and $1,594, handled the escapes, read 03/04/2026 as April 3, used the $100 amount due, and marked the deposit invoice unpaid. The looser modes had their own slips. Under json_object, gpt-oss-120b returned 50 line items for the 40-item invoice and renamed unit_price to unit on the escapes case. GPT-4.1 Mini marked the half-paid invoice as paid. Some models listed the credit note as a third line item and some left it out. We count both as right, because the schema does not say.

A better schema fixed what strict mode broke, on 5 of 6 OpenAI modelsHard invoice cases answered correctly, out of 8. Grey: strict, original schema. Light: json_object. Blue: strict, redesigned schema.GPT-6 Sol7 · strict, original8 · json_object8 · strict, redesignedGPT-6 Luna7 · strict, original8 · json_object8 · strict, redesignedGPT-6.1 Sol7 · strict, original8 · json_object8 · strict, redesignedGPT-5.4 Mini6 · strict, original6 · json_object8 · strict, redesignedGPT-4.1 Mini6 · strict, original6 · json_object8 · strict, redesignedgpt-oss-120b6 · strict, original6 · json_object5 · strict, redesignedScale: width = cases × 66 px (8 = 528 px). 2026-10-03 via OpenRouter. Rubric in the method section.
The blue bars are the same models, the same invoices and still strict mode. Only the schema changed.

The schema change that fixed it

Each wrong answer above had the same cause. The true value had no legal place in the schema. So we changed three fields and kept strict mode:

"currency":      {"type": "string", "enum": ["USD","EUR","GBP","JPY","OTHER"]},
"currency_code": {"type": "string", "description": "ISO 4217 code exactly as on the document"},
"date":          {"type": ["string","null"], "description": "YYYY-MM-DD, or null if the document shows no date"},
"qty":           {"type": "number"}

We also added one sentence to the prompt saying when to use OTHER and null. Then we re-ran the eight hard invoices. GPT-6 Sol, GPT-6 Luna, GPT-6.1 Sol, GPT-5.4 Mini and GPT-4.1 Mini all scored 8 out of 8. Every one wrote OTHER with CHF, null for the missing date, and 2.5 and 1.5 hours at $200 and $150. The structured-output feature did not change. The schema stopped forcing the model to choose between valid and true.

gpt-oss-120b and the whitespace runaway

gpt-oss-120b went the other way: 6 of 8 with the first schema, 5 of 8 with the redesigned one. On two invoices it produced spaces, tabs and newlines until it reached the 4000-token limit, ending with finish_reason: length and no closing brace. On a third it got the first line item right and then made up three more, including a quantity of 620 at a unit price of −19. It is an open-weight model served by third-party providers, which may implement constrained decoding differently from OpenAI's own API. Whatever the cause, it is a reason to keep a sane max_tokens and to treat length as a failure even when the response is valid JSON so far. DeepSeek V4.1 Flash, which we tested as a non-OpenAI control, also hit the token limit on two invoices with the new schema and returned no JSON at all.

Rules for OpenAI schemas

  1. Give every enum an escape value. OTHER plus a free-text field beats a forced guess.
  2. Make optional facts nullable. ["string","null"] is allowed in strict mode. Without it, the model must invent something or write an empty string.
  3. Type fields for the data, not for your database. Quantities can be fractional. Cast after validating, not before.
  4. Check the arithmetic yourself. Sum the line items against the total. Our h03 failures pass every schema check.
  5. Treat finish_reason: length as failure. Strict mode can produce valid-looking output until the budget runs out.
  6. Strip code fences when you do not use response_format. GPT-4.1 Mini added one to every prompt-only answer.

For the full list of which request parameters each model accepts, see our Chat Completions parameter test. Structured outputs and tools share the same schema dialect, so our function calling guide applies here too.

How we tested

All calls ran on 2026-10-03 through OpenRouter's OpenAI-compatible endpoint, not through the DataLLM Lab gateway, with max_tokens 4000 and one call per invoice per mode. The invoices are ones we wrote for this test: 10 clean and 8 hard, each with a known answer. A response counted as schema-valid if it parsed as bare JSON and passed our validator: required keys, no extra keys, enum, types, and integer qty. It counted as right under this rubric. h01: currency CHF or null, total 610. h02: date empty or null. h03: total 725 and line items summing to 725. h04: 40 items and 1,594. h05: 3 items and schema-valid. h06: 2026-04-03. h07: total 100. h08: paid false and total 2,400. For the redesigned schema, h01 required OTHER with CHF. Across all 13 models we tested, including the Gemini, Claude and DeepSeek runs, the clean set cost $0.7774, the hard set $1.1124 and the redesigned-schema run $0.3062, as API-reported usage.cost.

What this cannot tell you

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Primary sources, checked October 3, 2026: OpenRouter structured outputs. Dated measurements above may differ from the current documentation.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.