All WordPress HTML Templates Forms & Webhooks AI & Tools
AI & Tools

Structured Output From LLMs: JSON Schema, Validation and Repair

A practitioner’s guide to reliable LLM JSON output: the strict JSON Schema subset, function calling json, three-layer validation, repair turns and grammars.

A real working moment illustrating the theme of an article about llm json output. Wide 16:9 banner, one strong focal point, magazine editorial quality, authentic and unstaged.

A classification endpoint we built for a logistics client ran fine for six weeks. Then one afternoon it started throwing parse errors at a rate of about one in forty requests. Nothing had changed on our side. The model had been updated, and it had started wrapping its output in a markdown fence roughly 2% of the time.

That was 2023. We’d been doing what everyone did back then: asking nicely for JSON, then praying. The fix list ran to strip-the-fence, balance-the-braces, retry-on-failure. It worked, mostly, which is the worst kind of working.

Reliable llm json output is now a solved problem at the syntax layer, and almost entirely unsolved at the semantic layer. If you’re still writing brace-balancing code in 2026 you’re fixing the easy half. This is the version of the article written after shipping both halves.

Key Takeaways

  • Constrained decoding (OpenAI strict: true, Gemini responseSchema, XGrammar or Outlines on vLLM) makes malformed JSON structurally impossible. It does not make the contents correct, so schema validation is necessary but never sufficient.
  • The supported JSON Schema subset is narrower than you think: additionalProperties: false and every key in required are mandatory for strict mode, optional fields must be modelled as ["string","null"] unions, and most length or pattern constraints are either ignored or rejected depending on provider.
  • Field order in the schema changes output quality because generation is left to right. Put a short reasoning or evidence field before the decision field, never after it.
  • Validate in three layers: syntax, schema, then business rules (does this SKU exist, is this date within range). Feed layer-three failures back as a single repair turn with the exact validator message, and cap retries at one.
  • Track schema validation failure rate as a first-class metric. On production extraction work ours sits under 0.5%; a jump above 2% has always meant either a model version change or a prompt regression, never bad luck.

Syntax is solved. Semantics are not.

Constrained decoding works by masking the token distribution at every step so only tokens that can continue a valid parse of your grammar are sampled. If the schema says the next thing must be "status": followed by one of three enum values, the tokens for anything else get a probability of zero. There is no sampling path to invalid output. That’s a hard guarantee, not a nudge.

What people then assume is that a valid document is a correct document. It isn’t. A model forced into an enum of ["approved","rejected","needs_review"] will always return one of those three. When it has no idea, it returns the one that’s most statistically comfortable, usually the first or the most common in training data, with total syntactic confidence. We’ve seen a contract extraction pipeline return perfectly valid "jurisdiction": "Delaware" for documents that never mentioned a jurisdiction at all.

Constrained decoding removed a class of 500 errors and replaced it with a class of silent wrong answers. The second class is more expensive. Design accordingly: give the model an explicit escape hatch in every schema, which usually means a nullable field plus a confidence or not_found signal it’s allowed to use.

What the JSON Schema subset actually allows

Nobody reads this part until a deploy fails. Strict structured output on the major APIs accepts a subset of JSON Schema, and the subset is enforced at request time. An unsupported keyword is a 400, not a warning.

The rules that trip people up most often:

  • additionalProperties: false is required on every object, including nested ones. Omit it on one nested object three levels down and the whole request is rejected.
  • Every property must be listed in required. Optionality is expressed in the type, not by absence. So an optional string becomes "type": ["string", "null"].
  • anyOf is supported. allOf, oneOf and not generally are not. Discriminated unions work, but you build them with anyOf plus a literal const tag field.
  • $ref and recursive references work, including root recursion, which matters if you’re extracting tree-shaped data like nested document outlines.
  • String and numeric constraints (minLength, maxLength, pattern, minimum) have patchy support that has shifted across provider versions. Test them against the live API before you depend on them, and validate them again on your side regardless.

There are also size ceilings on total properties, nesting depth and the combined length of all property names and enum values. They’re generous enough that a sensible schema never hits them. If you are hitting them, the schema is the problem. A 200-field flat object is not an extraction schema, it’s a database table that wandered into the wrong room.

// Strict-mode schema for invoice extraction. Note the nullable unions,
// the const-tagged union, and additionalProperties on EVERY object.
const invoiceSchema = {
  name: "invoice_extraction",
  strict: true,
  schema: {
    type: "object",
    additionalProperties: false,
    // evidence comes first on purpose: it is generated before the decision
    required: ["evidence", "suppliername", "totalcents", "currency", "line_items", "confidence"],
    properties: {
      evidence: {
        type: "string",
        description: "Quote the exact text you used for the total and currency. Empty string if absent."
      },
      supplier_name: { type: ["string", "null"] },
      total_cents: { type: ["integer", "null"] },
      currency: { type: ["string", "null"], enum: ["USD", "EUR", "GBP", "INR", null] },
      line_items: {
        type: "array",
        items: {
          type: "object",
          additionalProperties: false,
          required: ["kind", "description", "amount_cents"],
          properties: {
            kind: { type: "string", enum: ["goods", "service", "tax", "discount"] },
            description: { type: "string" },
            amount_cents: { type: "integer" }
          }
        }
      },
      confidence: { type: "string", enum: ["high", "medium", "low"] }
    }
  }
};

Money as integer cents, not float. Currency as an enum including null. Both of those came from production bugs, not taste.

Function calling JSON versus a response format

These two mechanisms overlap enough that teams pick one by accident and then fight the consequences. The distinction is about control flow, not output quality.

Use a response format (response_format with a JSON schema, or Gemini’s responseSchema) when you know you want structured data back from this call. It’s a single deterministic shape: one turn, one object, done.

Use function calling json when the model needs to choose: which tool, whether to call one at all, or several in sequence. The schema lives on the tool’s inputschema, and the model’s arguments arrive as JSON you then execute against. On Anthropic’s API, forcing toolchoice to a specific tool is the idiomatic way to get guaranteed structured output, since it turns the tool into an output contract.

One trap: tool arguments on some providers are not under the same hard constraint as a strict response format. The schema is a strong hint plus server-side validation rather than token-level masking, depending on provider and model. Treat tool arguments as untrusted input, always. You’re about to execute them.

If you’re wiring tool calls into real systems, the dispatch layer matters as much as the schema. We put anything with side effects behind a queue rather than running it inline, for the same reasons we don’t process webhooks inline: retries, idempotency keys and the ability to replay a bad batch after you fix the prompt.

Design the schema for the model, not for your database

Generation is left to right. Every field you emit conditions the next one. That single fact drives most of what follows.

Put reasoning before conclusions

If decision comes first and rationale second, the rationale is a post-hoc justification of a token the model already committed to. Flip them and accuracy improves measurably. On a support-ticket triage schema we ran both orders against 500 hand-labelled tickets: evidence-first was right on 89% versus 81% for decision-first, same model, same prompt, same temperature. That’s free accuracy from reordering two keys.

Prefer enums over free strings, but keep them short

An enum of 8 categories is a gift to a constrained decoder. An enum of 400 product codes is a liability: it bloats the schema, eats input tokens, and the model will confidently pick a neighbour. Above roughly 50 values, do retrieval first and pass the shortlist in the prompt, or match the free-text output against your canonical list afterwards with a similarity threshold. That’s the point where a vector lookup earns its place.

Flatten, and version

Deep nesting costs tokens and confuses models. Four levels is usually a modelling failure. Put a schema_version in your logs (not in the schema itself, where it wastes tokens) so that when failure rates move you can tell which contract was live.

Validation in three layers

Layer one is syntax: did it parse. With constrained decoding this should be effectively 100%, and any failure means truncation, not malformation. Layer two is schema: does it match the contract. Layer three is the one everyone skips, and it’s where the value is: do the values mean anything.

import { z } from "zod";
import { generateObject } from "ai";

const Invoice = z.object({
  evidence: z.string(),
  supplier_name: z.string().nullable(),
  total_cents: z.number().int().nullable(),
  currency: z.enum(["USD", "EUR", "GBP", "INR"]).nullable(),
  line_items: z.array(z.object({
    kind: z.enum(["goods", "service", "tax", "discount"]),
    description: z.string(),
    amount_cents: z.number().int()
  })),
  confidence: z.enum(["high", "medium", "low"])
});

// Layer 3: business rules the schema cannot express
const InvoiceChecked = Invoice.superRefine((inv, ctx) => {
  const sum = inv.lineitems.reduce((a, l) => a + l.amountcents, 0);
  if (inv.totalcents !== null && Math.abs(sum - inv.totalcents) > 100) {
    ctx.addIssue({
      code: "custom",
      path: ["total_cents"],
      // this message is what we send back to the model on repair
      message: lineitems sum to ${sum} but totalcents is ${inv.total_cents}
    });
  }
  if (inv.total_cents !== null && inv.currency === null) {
    ctx.addIssue({ code: "custom", path: ["currency"], message: "currency required when total_cents is set" });
  }
  if (inv.confidence === "high" && inv.evidence.trim() === "") {
    ctx.addIssue({ code: "custom", path: ["evidence"], message: "confidence high requires quoted evidence" });
  }
});

const { object } = await generateObject({
  model: openai("gpt-4.1"),
  schema: Invoice,          // sent to the API as strict JSON Schema
  prompt: documentText,
  temperature: 0
});

const result = InvoiceChecked.safeParse(object);

The arithmetic check caught more real extraction errors on that project than every other validation combined. Cross-field consistency is the cheapest hallucination detector you will ever write, because a model that invented a number rarely invents a set of numbers that add up.

Repair, retry, and when to give up

Repair splits into two completely different problems that get discussed as one.

Syntactic repair is for unconstrained output: markdown fences, trailing commas, single quotes, unescaped newlines inside strings, a truncated tail. If you must handle this, use jsonrepair rather than your own regex, and only as a fallback behind the real parse. In 2026, needing syntactic repair on a provider that supports strict mode means one of two things: you’ve skipped strict mode, or you’ve hit max_tokens mid-object. The second one is not a repair problem. Raise the limit or shrink the schema.

Semantic repair is the useful kind. One extra turn, carrying the invalid object and the exact validator messages:

// one repair turn, then fail loudly. Loops here burn money and hide bugs.
if (!result.success) {
  const issues = result.error.issues
    .map(i => ${i.path.join(".")}: ${i.message})
    .join("\n");

  const retry = await generateObject({
    model: openai("gpt-4.1"),
    schema: Invoice,
    temperature: 0,
    prompt: [
      documentText,
      Your previous answer failed validation:\n${issues},
      Previous answer:\n${JSON.stringify(object)},
      Fix only the listed problems. If the source document does not contain the value, use null and set confidence to "low".
    ].join("\n\n")
  });
  // re-validate; if it fails again, route to human review, do not loop
}

Cap it at one retry. Our measured second-attempt success rate on business-rule failures is around 60 to 70%; the third attempt adds almost nothing and roughly doubles the latency budget for the whole request. The explicit instruction to use null matters more than it looks. Without it, the model “fixes” the arithmetic by inventing a line item that balances the total. It will do this. Confidently.

Anything that fails twice goes to a human queue with the raw document, the two attempts and the validator output attached. That queue is not a failure of the system, it’s part of the system, and it’s the same principle as designing the states where things go wrong rather than pretending they won’t.

Local models, grammars and the latency bill

Self-hosting changes the mechanics but not the approach. On vLLM you get guided decoding with a choice of backend, usually XGrammar or Outlines, selected per request or at server start. On llama.cpp you can go lower level with GBNF grammars, which is the right tool when your output isn’t JSON at all (a DSL, a SQL subset, a specific date format).

# vLLM: OpenAI-compatible, guided JSON via the structured outputs backend
vllm serve Qwen/Qwen3-8B \
  --guided-decoding-backend xgrammar \
  --max-model-len 16384

curl http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen/Qwen3-8B",
    "messages": [{"role":"user","content":"Extract: ACME Ltd, 2 units, 49.99 USD"}],
    "temperature": 0,
    "response_format": {
      "type": "json_schema",
      "json_schema": {"name":"line","strict":true,"schema":{
        "type":"object","additionalProperties":false,
        "required":["supplier","qty","amount_cents","currency"],
        "properties":{
          "supplier":{"type":"string"},
          "qty":{"type":"integer"},
          "amount_cents":{"type":"integer"},
          "currency":{"type":"string","enum":["USD","EUR","GBP"]}
        }}}
    }
  }'

Two latency facts worth budgeting for. First, a new schema has to be compiled into a token mask automaton, and that compile shows up on the first request that uses it. Hosted APIs cache the compiled artifact per schema, so the first call after a schema change is slower and the rest are not. Generate schemas dynamically per request and you pay that cost every single time. That’s one of the few genuinely bad ideas in this area.

Second, per-token masking overhead on modern backends is small enough to disappear into GPU time for typical schemas. Where it stops being small is wide enum sets and heavy recursion, because the mask has to be recomputed over a large candidate space. If your tokens per second drops noticeably after adding a schema, look at your enums first.

Local inference also means you own the evaluation harness. Fixed test set, versioned prompts, failure rate tracked per schema version. The discipline is the same as any other AI feature going from prototype to production: without a measured baseline you cannot tell a model upgrade from a regression.

Streaming partial objects

Users will wait 400ms. They won’t happily wait 9 seconds staring at a spinner while a 30-field object assembles. Streaming structured output means parsing incomplete JSON, which needs a partial parser (the Vercel AI SDK’s streamObject, Instructor’s partial mode, or partial-json directly) rather than JSON.parse in a try block.

The rule we’ve settled on: only render a field once its value is syntactically complete. A string that’s still streaming can flicker through states that look like answers and aren’t, and a half-written number is a wrong number. For arrays, render completed elements and show a skeleton row for the one in flight. Never run layer-three validation on a partial object. The arithmetic won’t balance until the last line item lands, so you’ll flash errors at users for no reason.

Frequently Asked Questions

Does strict structured output make models less accurate?

Not in our testing, as long as the schema leaves room to reason. The early reports of quality loss came from schemas that forced an answer token first with no space for intermediate thinking. Add an evidence or reasoning string as the first property and the gap closes; on our triage benchmark it reversed into an 8 point gain over free-text JSON parsing.

Should I use a schema library or hand-write JSON Schema?

Use a library. Pydantic in Python, Zod in TypeScript, both of which emit JSON Schema and give you runtime validation from the same definition. Hand-written JSON Schema drifts from your validator within weeks, and then you get objects that pass the API’s contract and fail yours. The one exception is when you need an unusual construct the converter mangles, in which case keep the hand-written schema next to the validator and test that they agree.

How do I handle a model that refuses to answer?

Treat refusals as a separate branch, not a parse failure. OpenAI returns a dedicated refusal field on the message when a safety system declines, and your code should check it before touching the parsed content. On other providers you may get a valid object full of nulls or an empty tool call, so make your pipeline tolerate a well-formed empty answer without retrying forever.

What failure rate should I expect in production?

With strict mode plus sensible schema design, syntax and schema failures should be effectively zero and business-rule failures depend entirely on your input quality. Clean digital documents land under 1% for us; photographed receipts run 5 to 15%. Set the alert threshold relative to your own measured baseline, and investigate any step change, because they almost always trace to a model version rollout or an edited prompt.

Can I trust the confidence field the model returns?

Partially. Self-reported confidence correlates with correctness well enough to route work, and badly enough that you should not use it as a gate on anything irreversible. Pair it with a hard check: cross-field arithmetic, a lookup against your own data, or a required quoted span from the source. Those signals are objective in a way that a self-assessment never is, which is also the core of any workable guardrail setup for customer-facing AI.

If you’re starting today, the order is: strict schema from a validator library, reasoning field first, business rules in layer three, one repair turn, human queue behind that. Then instrument the validation failure rate and leave it on a dashboard where you’ll see it.

The decision worth making deliberately is how much you trust the parsed object. The honest answer is that it’s as trustworthy as your layer-three checks, and nothing about the schema changes that. Write the cross-field check first. It will tell you more about your pipeline in a week than any amount of prompt tuning.

constrained decoding grammar vllm function calling json validation json repair llm retry strategy llm json output openai strict mode json schema streaming partial json llm structured output json schema zod pydantic llm schema validation