OpenAI Structured Outputs: Valid JSON Is Not the Same as a Right Answer (Tested)
Test scope: synthetic invoice fixtures through OpenRouter, not a direct-vendor endpoint test. OpenAI structured outputs promise JSON that matches your schema, and on our test they delivered. Under strict: true all six OpenAI models we tried returned schema-valid JSON on all 8 of our deliberately awkward invoices. What the schema cannot promise is that the values are true. When an invoice was in Swiss francs and the schema only allowed USD, EUR, GBP or JPY, all six models picked one of the four. The JSON was valid, and the currency was wrong. When we redesigned the schema so that the true answer had somewhere to go, five of the six scored 8 out of 8.
Structured outputs solve a real problem. Before them, a model could return a broken bracket, a trailing comment or a field you never asked for, and your parser would crash at 3am. Strict mode removes that whole class of failure. This page is about the failure it does not remove. We ran the same invoice extraction three ways on six OpenAI models on 2026-10-03, scored the answers as well as the shape, and then tested a fix. If you use Claude, our Claude structured output guide covers the Anthropic side. The Gemini results are in a separate Gemini test.
Three ways to ask for JSON
json_schemawithstrict: true: OpenAI structured outputs. The response must match your schema, and generation is constrained so that it does.json_object: the older JSON mode. The output is meant to be valid JSON, but which keys appear is up to the model and your prompt.- Prompt only: no
response_format, just “return only the JSON object”.
Our schema asked for vendor, date (YYYY-MM-DD), currency (an enum of USD, EUR, GBP, JPY), total, paid (boolean), and line_items, each with description, qty (integer) and unit_price. additionalProperties was false everywhere. It is a typical first-draft schema. That is the point.
Clean invoices: everyone passes
We first ran 10 ordinary invoices: different languages and date formats, USD, EUR, GBP and JPY, paid and unpaid. All six models returned schema-valid JSON with the right total, currency, date, paid flag and item count on all 10, in all three modes. One thing differed. In prompt-only mode, GPT-4.1 Mini wrapped its JSON in a markdown code fence on all 10. A plain json.loads fails on every one of those. The other five returned bare JSON. On clean input, strict mode's only advantage is that you never have to strip a fence.
Eight awkward invoices
Then we wrote eight invoices that break the first-draft schema on purpose. Each one tests a single thing:
| Case | The catch | Right answer |
|---|---|---|
| h01 | Hotel bill in CHF, which the enum does not allow | CHF, or admit it does not fit |
| h02 | No date anywhere on the document | Empty or null, not an invented date |
| h03 | 2.5 and 1.5 consulting hours, but qty must be an integer | Line items that still add up to the $725 total |
| h04 | 40 line items | 40 items, total $1,594 |
| h05 | Quotes, backslashes and a tab inside descriptions | 3 items, valid JSON |
| h06 | 03/04/2026 from a London vendor | 2026-04-03 |
| h07 | A $20 credit note, so the amount due is not the item sum | Total $100 |
| h08 | 50% deposit paid, balance outstanding | paid false, total $2,400 |
| Model | Strict: schema-valid | Strict: right | json_object: right | Prompt only: right |
|---|---|---|---|---|
| GPT-6 Sol | 8/8 | 7/8 | 8/8 | 8/8 |
| GPT-6 Luna | 8/8 | 7/8 | 8/8 | 8/8 |
| GPT-6.1 Sol | 8/8 | 7/8 | 8/8 | 8/8 |
| GPT-5.4 Mini | 8/8 | 6/8 | 6/8 | 6/8 |
| GPT-4.1 Mini | 8/8 | 6/8 | 6/8 | 7/8 |
| gpt-oss-120b | 8/8 | 6/8 | 6/8 | 8/8 |
Read it across. Strict mode had the best schema-validity of the three modes and, for the GPT-6 models, the worst correctness. The looser modes were right more often because the model could write "currency": "CHF" or "qty": 2.5. That broke the schema and kept the truth. A validator catches a schema break. Nothing catches a valid lie.
What strict mode did with each one
h01, the franc. Under strict mode GPT-6 Sol, GPT-5.4 Mini, GPT-4.1 Mini and gpt-oss-120b wrote USD. GPT-6 Luna and GPT-6.1 Sol wrote EUR. The total stayed at 610 on all six, so the amount was right and the currency was wrong. Without the schema, most models wrote CHF or null. GPT-5.4 Mini was the exception: it picked EUR under json_object and USD with the prompt alone, even though nothing forced it to.
h02, the missing date. All six returned an empty string under strict mode. That is the honest answer the schema allowed, since it required a string. Without the schema, all six returned null. No OpenAI model made up a date. Three Gemini models did, as the Gemini test shows.
h03, the fractional hours. This one splits the field. GPT-6 Sol, Luna and 6.1 Sol set each qty to 1 and moved the money into unit_price, $500 and $225. The line items still sum to $725. The others rounded the hours and kept the hourly rate. GPT-5.4 Mini wrote 3 and 2 hours, so its items sum to $900. GPT-4.1 Mini wrote 2 and 2, which comes to $700. gpt-oss-120b wrote 2 and 1, which comes to $550. All of those are schema-valid. Each reports a total of $725 next to line items that do not add up to it.
h04 to h08. Under strict mode all six got 40 items and $1,594, handled the escapes, read 03/04/2026 as April 3, used the $100 amount due, and marked the deposit invoice unpaid. The looser modes had their own slips. Under json_object, gpt-oss-120b returned 50 line items for the 40-item invoice and renamed unit_price to unit on the escapes case. GPT-4.1 Mini marked the half-paid invoice as paid. Some models listed the credit note as a third line item and some left it out. We count both as right, because the schema does not say.
The schema change that fixed it
Each wrong answer above had the same cause. The true value had no legal place in the schema. So we changed three fields and kept strict mode:
"currency": {"type": "string", "enum": ["USD","EUR","GBP","JPY","OTHER"]},
"currency_code": {"type": "string", "description": "ISO 4217 code exactly as on the document"},
"date": {"type": ["string","null"], "description": "YYYY-MM-DD, or null if the document shows no date"},
"qty": {"type": "number"}
We also added one sentence to the prompt saying when to use OTHER and null. Then we re-ran the eight hard invoices. GPT-6 Sol, GPT-6 Luna, GPT-6.1 Sol, GPT-5.4 Mini and GPT-4.1 Mini all scored 8 out of 8. Every one wrote OTHER with CHF, null for the missing date, and 2.5 and 1.5 hours at $200 and $150. The structured-output feature did not change. The schema stopped forcing the model to choose between valid and true.
gpt-oss-120b and the whitespace runaway
gpt-oss-120b went the other way: 6 of 8 with the first schema, 5 of 8 with the redesigned one. On two invoices it produced spaces, tabs and newlines until it reached the 4000-token limit, ending with finish_reason: length and no closing brace. On a third it got the first line item right and then made up three more, including a quantity of 620 at a unit price of −19. It is an open-weight model served by third-party providers, which may implement constrained decoding differently from OpenAI's own API. Whatever the cause, it is a reason to keep a sane max_tokens and to treat length as a failure even when the response is valid JSON so far. DeepSeek V4.1 Flash, which we tested as a non-OpenAI control, also hit the token limit on two invoices with the new schema and returned no JSON at all.
Rules for OpenAI schemas
- Give every enum an escape value.
OTHERplus a free-text field beats a forced guess. - Make optional facts nullable.
["string","null"]is allowed in strict mode. Without it, the model must invent something or write an empty string. - Type fields for the data, not for your database. Quantities can be fractional. Cast after validating, not before.
- Check the arithmetic yourself. Sum the line items against the total. Our h03 failures pass every schema check.
- Treat
finish_reason: lengthas failure. Strict mode can produce valid-looking output until the budget runs out. - Strip code fences when you do not use
response_format. GPT-4.1 Mini added one to every prompt-only answer.
For the full list of which request parameters each model accepts, see our Chat Completions parameter test. Structured outputs and tools share the same schema dialect, so our function calling guide applies here too.
How we tested
All calls ran on 2026-10-03 through OpenRouter's OpenAI-compatible endpoint, not through the DataLLM Lab gateway, with max_tokens 4000 and one call per invoice per mode. The invoices are ones we wrote for this test: 10 clean and 8 hard, each with a known answer. A response counted as schema-valid if it parsed as bare JSON and passed our validator: required keys, no extra keys, enum, types, and integer qty. It counted as right under this rubric. h01: currency CHF or null, total 610. h02: date empty or null. h03: total 725 and line items summing to 725. h04: 40 items and 1,594. h05: 3 items and schema-valid. h06: 2026-04-03. h07: total 100. h08: paid false and total 2,400. For the redesigned schema, h01 required OTHER with CHF. Across all 13 models we tested, including the Gemini, Claude and DeepSeek runs, the clean set cost $0.7774, the hard set $1.1124 and the redesigned-schema run $0.3062, as API-reported usage.cost.
What this cannot tell you
- OpenAI's direct API. We used OpenRouter, which forwards
response_formatto OpenAI for the GPT models. gpt-oss-120b went to third-party providers. - One run per cell. 7 of 8 against 8 of 8 is one invoice. The schema effect is large and consistent across models. The gaps between models are not.
- Invoices we wrote. They are short and built to break a specific schema. Scanned documents and long contracts fail in other ways.
- Other schemas. A schema that already allows nullables and escape values would have done better in the first round. That is our finding, not a caveat.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
Primary sources, checked October 3, 2026: OpenRouter structured outputs. Dated measurements above may differ from the current documentation.
DataLLM Lab