Developer Guide

Gemini Structured Output: Valid JSON, and Three Models Invented a Date (Tested)

Test scope: synthetic invoice fixtures through OpenRouter, not a direct-vendor endpoint test. We tested Gemini structured output on 2026-10-03 with five Gemini models and 18 invoices. Gemini 3.8 Flash, 3.7 Flash, 3.5 Flash Lite, 3.1 Pro Preview and 2.5 Flash each ran three ways: a strict JSON schema, JSON mode and prompt only. Shape was rarely the problem. The answers were. Given an invoice with no date on it and a schema that required one, three of the five wrote a date anyway: 2024-01-01, 2023-01-01 and 2023-10-24. No OpenAI model we tested did that. Once the schema allowed null, all five Gemini models scored 8 out of 8.

DataLLM Lab article cover: Gemini Structured Output: Valid JSON, and Three Models Invented a Date (Tested)

This is the Gemini half of a test we ran on 13 models. The method and the invoices are the same as in our OpenAI structured outputs test, so you can compare the two families line by line. This page covers what is specific to Gemini: invented values, code fences, one truncated answer, and the bill.

What we tested, and through what

We called all five models through OpenRouter's OpenAI-compatible endpoint, using response_format. OpenRouter sent the calls to Google's own endpoints, which appeared as Google and Google AI Studio in the responses. We did not call the native Gemini API with its own response schema field. If you use Google's SDK directly, the constraint mechanism may differ. Treat this as a test of Gemini models behind the OpenAI-style interface, which is how most multi-model apps call them.

The schema was a deliberately ordinary first draft. It asked for vendor, an ISO date, a currency from USD, EUR, GBP or JPY, a total, a paid flag, and line items with an integer quantity and a unit price. No extra keys were allowed.

Clean invoices and the code-fence problem

On 10 ordinary invoices all five models were right on every field, in every mode. The difference showed up in the raw text. With the prompt alone, no response_format, several Gemini models wrapped the JSON in a markdown code fence:

ModelFenced answers, clean set (of 10)Fenced answers, hard set (of 8)
Gemini 2.5 Flash108
Gemini 3.7 Flash46
Gemini 3.5 Flash Lite35
Gemini 3.8 Flash12
Gemini 3.1 Pro Preview10

A fenced answer is valid JSON inside a fence, and json.loads rejects it. With json_object or a strict schema, none of the five fenced anything. If you call Gemini without response_format, strip fences before parsing. Gemini 2.5 Flash fenced every one of its 18 prompt-only answers.

Eight awkward invoices

The hard set has eight invoices, each built to break the first-draft schema in one way. One is in Swiss francs. One has no date. One bills 2.5 and 1.5 hours. One has 40 line items. One has quotes and backslashes inside descriptions. One shows 03/04/2026 from a London vendor. One has a credit note, and one has a 50% deposit. The OpenAI article lists the cases and our scoring rubric in full. Here are the Gemini results:

ModelStrict: schema-validStrict: rightjson_object: rightPrompt only: right
Gemini 3.8 Flash8/86/88/88/8
Gemini 3.7 Flash8/86/88/88/8
Gemini 3.5 Flash Lite8/85/87/87/8
Gemini 3.1 Pro Preview7/85/88/87/8
Gemini 2.5 Flash8/85/88/87/8

The pattern matches the OpenAI models, and it is a bit stronger. Strict mode had the best schema-validity and the worst correctness on all five. JSON mode, which let the model write values outside the schema, was right 8 of 8 on four of them.

All five Gemini models reached 8 of 8 once the schema could hold the truthHard invoice cases answered correctly, out of 8. Grey: strict, original schema. Light: json_object. Blue: strict, redesigned schema.Gemini 3.8 Flash6 · strict, original8 · json_object8 · strict, redesignedGemini 3.7 Flash6 · strict, original8 · json_object8 · strict, redesignedGemini 3.5 Flash Lite5 · strict, original7 · json_object8 · strict, redesignedGemini 3.1 Pro Preview5 · strict, original8 · json_object8 · strict, redesignedGemini 2.5 Flash5 · strict, original8 · json_object8 · strict, redesignedScale: width = cases × 66 px (8 = 528 px). 2026-10-03 via OpenRouter. Rubric in the method section.
Same models, same eight invoices, still strict mode. Only the schema changed between the grey and the blue bars.

The invented dates

Case h02 is a copywriting invoice with no date anywhere. The schema required date as a string. Here is what each model wrote under strict mode:

Modeldate, strict schemadate, json_object
Gemini 3.8 Flash""null
Gemini 3.7 Flash""null
Gemini 3.5 Flash Lite2023-10-24null
Gemini 3.1 Pro Preview2023-01-01null
Gemini 2.5 Flash2024-01-01null

All five knew there was no date. Each wrote null as soon as the schema allowed it. Under the strict schema, three filled the field with a real-looking ISO date. That date would pass any validator and could end up in your ledger. All six OpenAI models, Claude Sonnet 5.5 and DeepSeek V4.1 Flash wrote an empty string in the same position. Gemini 3.1 Pro Preview, which ran up the largest bill of the 13 models we tested (see the cost table below), is one of the three.

Currency, hours and a truncated answer

The franc. All five swapped CHF for an allowed currency under strict mode. Gemini 3.8 Flash, 3.7 Flash and 3.1 Pro Preview wrote EUR, and 3.5 Flash Lite and 2.5 Flash wrote USD. Without a schema, 2.5 Flash still wrote GBP with the prompt alone. The others wrote CHF or null.

The fractional hours. Under the integer constraint only Gemini 3.1 Pro Preview kept the money right. It used a quantity of 1 at $500 and $225, so the items sum to the $725 total. Gemini 3.7 Flash, 3.5 Flash Lite and 2.5 Flash rounded to 2 and 1 hours, which sums to $550. Gemini 3.8 Flash rounded up to 3 and 2, which sums to $900. Gemini 3.5 Flash Lite rounded to 3 and 2 even in JSON mode and prompt-only mode, where it could have written 2.5. The other four kept 2.5 and 1.5 whenever the schema allowed it.

The 40-item invoice. Gemini 3.1 Pro Preview returned JSON that stopped partway through the array, both under strict mode and with the prompt alone. We re-ran the strict call four times, twice at max_tokens 4000 and twice at 16000. All four completed with finish_reason: stop, used between 795 and 2,195 reasoning tokens, and parsed. We could not reproduce the cut-off, so we report it as intermittent. Because reasoning tokens count against the same limit, a 4000-token cap leaves less room than it seems on long outputs.

The schema that took all five to 8 of 8

Each wrong answer had one cause: the truth had nowhere legal to go. We kept strict mode and changed the schema. Currency gained an OTHER value plus a free currency_code field. Date became ["string","null"]. Quantity became a number. One prompt sentence said when to use OTHER and null. On the same eight invoices, all five Gemini models scored 8 out of 8. Each wrote OTHER with CHF, null for the missing date, and 2.5 and 1.5 hours. Of the families we tested with more than one model, Gemini was the only one where every model reached 8 of 8. gpt-oss-120b kept the OpenAI group from doing the same. The exact schema fragment is in the OpenAI test.

What it cost

Here is the API-reported usage.cost per model across all three runs: 10 clean invoices in three modes, 8 hard invoices in three modes, and 8 with the redesigned schema.

ModelClean setHard setRedesigned schema
Gemini 3.1 Pro Preview$0.3897$0.5082$0.1296
Gemini 3.8 Flash$0.0880$0.1354$0.0294
Gemini 3.7 Flash$0.0772$0.0994$0.0207
Gemini 3.5 Flash Lite$0.0143$0.0218$0.0081
Gemini 2.5 Flash$0.0126$0.0189$0.0052

Gemini 3.1 Pro Preview accounted for $0.3897 of the $0.7774 clean-set total across all 13 models, half the bill for one model out of thirteen. On extraction this simple, it was not more accurate than Gemini 2.5 Flash. Both scored 8 of 8 with the redesigned schema, and 2.5 Flash cost about a twenty-fifth as much ($0.0052 against $0.1296). For extraction work, start with a Flash tier and move up only if your own error log tells you to. Our Gemini API pricing guide has the per-token rates, and our Gemini 3.8 Flash review has its coding results.

How we tested

All calls ran on 2026-10-03 through OpenRouter, not through the DataLLM Lab gateway, with max_tokens 4000 and one call per invoice per mode. We wrote the 18 invoices for this test, and each has a known answer. A response counted as schema-valid if it parsed as bare JSON and passed our validator. It counted as right under the rubric listed in the OpenAI article: no invented date, the right currency or an honest null, line items that add up to the total, and the right amount due and paid status. The four Gemini 3.1 Pro Preview re-runs on the 40-item case were made the same day.

What this cannot tell you

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Primary sources, checked October 3, 2026: Gemini structured output documentation. Dated measurements above may differ from the current documentation.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.