AI Structured Outputs Explained: JSON Mode, Schemas, Function Calling, and the Reliability Gap
AI CORE · CONCEPT GUIDE

AI Structured Outputs Explained: JSON Mode, Schemas, Function Calling, and the Reliability Gap

AI structured outputs make language-model responses easier for software to parse, display, route, and validate. They solve the shape of an answer—but not automatically its truth. This guide explains the difference, shows where JSON mode and function calling fit, and gives you a practical validation model for real applications.

Updated September 202622 min read
A human designer watching free-form notes pass through a transparent schema engine and emerge as organized structured data blocks
SJ

Written by

Peter M · Singularity Journey

Research-backed guidance for building practical skills around emerging AI systems.

Reviewed for clarity, source quality, and practical usefulness for Singularity Journey readers.

AI Structured Outputs: The Quick Answer

AI structured outputs are model responses constrained to a defined machine-readable shape, commonly a JSON Schema. Instead of receiving an unpredictable paragraph, an application can request fields such as category, priority, summary, and needs_review with specified data types and allowed values. This makes downstream parsing far more dependable. It does not prove that the values inside those fields are factually correct, supported by the input, authorized, or safe to act on.

The cleanest mental model is: a schema is a mold, not a fact-checker. It can require the model to produce a cube instead of a blob. It cannot guarantee that the cube contains the right information. A response can be valid JSON, match every field in your schema, and still contain an invented invoice number, an unsupported diagnosis, a wrong customer status, or a confidently misclassified request.

That distinction matters because structured output often sits at the boundary between probabilistic AI and deterministic software. Once a response looks like ordinary application data, developers may be tempted to trust it like data from a database. The safer approach is to treat it as a typed proposal that must pass progressively stronger checks before it becomes a record, decision, tool call, or external action.

The practical rule: use structure to make AI output inspectable and testable. Then use application code, source evidence, business rules, permissions, and human review to decide whether the structured result deserves trust.

What Are AI Structured Outputs?

A language model naturally generates a sequence of tokens. Left unconstrained, it can answer with prose, a list, Markdown, a code block, partial JSON, or a mixture of all five. Humans can often interpret those variations. Software cannot safely assume that the field it needs will appear in the right place, use the right type, or even exist.

Structured output adds a contract for the response shape. The contract might say that the result must be an object with a required sentiment field chosen from three values, a numerical confidence field, an array of evidence strings, and a Boolean needs_human_review field. Modern model APIs can use a supported subset of JSON Schema to guide or constrain generation so the result is syntactically valid and adheres to the requested structure.

That makes structured output useful wherever an AI response must cross into conventional software: filling a form, rendering a UI, classifying support tickets, extracting line items, generating a workflow plan, passing data to another agent, or storing a result for later analysis. It replaces fragile instructions such as “return only JSON and do not add commentary” with an explicit interface.

The key word is interface. Structured output does not turn a language model into a database query engine, rules engine, or source of record. It gives those systems a predictable way to receive the model’s proposal. The receiving application still owns validation and consequences.

A small example

Imagine a customer email that says, “My order arrived damaged and I need a replacement before Friday.” A free-form assistant might write a sympathetic paragraph. A structured-output classifier could return an object like this:

{
  "intent": "replacement_request",
  "urgency": "high",
  "order_id": null,
  "requested_deadline": "Friday",
  "needs_human_review": true,
  "reason": "No order identifier was present in the message."
}

The shape helps the application route the request. Yet the application must still verify that “Friday” can be converted to a real date, that the customer is entitled to a replacement, that an order can be identified, and that the model did not infer details absent from the email. Good structure creates a reliable starting point for those checks.

Raw Text, JSON Mode, Structured Outputs, and Function Calling

These terms are often blended together because they can all involve JSON. They solve different problems. The official OpenAI structured outputs guide distinguishes schema-shaped responses from function calling, and the Gemini documentation makes the same primary-use distinction: structured outputs format a final answer, while function calling helps a model request an intermediate action or data lookup.

ModeWhat it guaranteesBest useWhat it does not guarantee
Free-form textNo machine-readable contract beyond normal language generationExplanations, drafting, brainstorming, conversational answersStable fields, parseability, consistent formatting, factual accuracy
Prompted JSONNothing formal; the prompt merely asks for JSONExperiments and models without a native structured modeValid JSON, schema adherence, no extra prose, correct values
JSON modeValid JSON at the syntax levelFlexible objects where exact schema enforcement is unavailable or unnecessaryRequired keys, exact types, allowed values, semantic correctness
Structured outputValid JSON that adheres to a supported schemaExtraction, classification, UI data, workflow state, agent handoffsTruth, evidence, authorization, business-rule validity, completeness of the source
Function or tool callingA structured request naming a tool and proposed argumentsRetrieving data or asking the application to perform an actionThat the tool should run, the arguments are safe, or the action is authorized
Isometric diagram showing one AI model branching into free-form text, JSON, schema-constrained data, and a tool-call handoff

One model can produce several kinds of output. Choose the mode according to what the receiving system must do next.

Why JSON mode is not the same as a schema

Valid JSON only means the braces, quotes, commas, arrays, and primitive values can be parsed. A model could return {"banana":42} when your application expected a support-ticket classification. The document is valid JSON and useless to the workflow. Schema adherence adds field names, types, required properties, enums, nesting rules, and other supported constraints.

Why function calling is not merely structured extraction

A function call is a proposal to cross an execution boundary. The model might request get_order_status with an order identifier, or cancel_subscription with a customer ID. The application or tool runtime executes the operation and returns a result. Structured output, by contrast, can simply format the model’s final answer for a UI or storage layer. Some APIs use the same schema machinery for both, but the consequence is different: one shapes data; the other may initiate behavior.

If you need a deeper explanation of agents that can plan and use external capabilities, read AI agent tool use explained and AI agents vs. agentic AI. The important connection is that every tool call is structured, but not every structured response is a tool call.

How Schema-Constrained AI Output Works

At a practical level, your application sends three things to the model service: the task, the source material or context, and a schema describing the desired response. The provider translates that schema into guidance and, on supported models, constraints over which tokens can legally appear as the response is generated.

  1. Define the job. State what should be extracted, classified, summarized, or decided. A schema cannot rescue an ambiguous task definition.
  2. Define the shape. Specify objects, arrays, strings, numbers, Booleans, nullability, required fields, and enums. Keep the schema as small as the downstream decision allows.
  3. Generate under constraints. The model produces tokens while the provider prevents disallowed structural paths. Exact behavior and the supported JSON Schema subset vary by provider and model.
  4. Parse and validate locally. The application deserializes the result into a typed object and re-validates it. Native schema adherence is not a reason to remove local checks.
  5. Apply domain checks. Verify dates, identifiers, totals, relationships, source evidence, policy rules, and any invariants the schema cannot express.
  6. Choose the consequence. Accept, reject, retry, route to review, or request missing information. High-impact outputs should not silently fall through to action.

Constrained generation can be more reliable than asking the model politely to follow a format because invalid structural tokens are blocked rather than merely discouraged. However, providers generally support only part of the full JSON Schema standard, and supported features change. Google documents that its structured mode supports a subset and warns that very large or deeply nested schemas may be rejected. Amazon Bedrock’s documentation similarly lists supported and unsupported schema features for its runtime.

What schema descriptions really do

A field description is not decoration. It is part of the model’s instruction context. Compare a vague field such as status with a precise one: “The customer-visible fulfillment state supported by the cited order record; never infer from tone.” The second description narrows the intended meaning and tells the model which evidence matters.

Even then, descriptions remain instructions to a probabilistic system. If your application needs a hard rule—“refund amount must not exceed captured payment”—implement it in deterministic code against authoritative records. Use the schema to transport the candidate value, not to enforce the financial policy.

The Reliability Gap: Valid Shape Is Not Valid Meaning

The most dangerous structured-output mistake is believing that a parseable, schema-valid object is automatically trustworthy. It looks like normal application data. It passes type checks. It may even contain a confidence score. None of those properties prove that the model read the source correctly or that the source supported the conclusion.

The Gemini documentation states this directly: syntactically correct JSON does not guarantee semantically correct values, so applications should validate final outputs and handle schema-compliant but incorrect results. This is not a minor caveat. It defines the boundary of what structured output can solve.

Syntactic validityThe output is parseable JSON. Braces close, strings are quoted, and the parser can load the document.
Structural validityThe output matches the schema: required fields exist, types fit, and allowed values are respected.
Semantic validityThe values accurately represent the source and mean what the application believes they mean.
Operational validityThe result is permitted, safe, timely, and appropriate for the action the system may take.
Evidence validityClaims can be traced to authoritative source material rather than model inference or fabrication.
Decision validityThe result satisfies business rules, risk thresholds, exception handling, and review requirements.

Five examples of perfectly shaped wrong answers

  • Invented identifier: the schema requires an invoice ID, so the model supplies a plausible-looking one even though the document is unreadable.
  • False enum certainty: the allowed values are approved or denied, but the evidence is incomplete. Without an unknown option, the model must choose a false certainty.
  • Wrong date interpretation: a valid ISO date is returned, but “next Friday” was resolved against the wrong timezone or reference date.
  • Correct type, wrong unit: a numeric amount is extracted as 1,200 while the source meant 1,200 cents, not 1,200 dollars.
  • Unsupported confidence: the schema asks for a number from zero to one, so the model produces 0.94. The number looks scientific but may not be calibrated against real outcomes.
Schema pressure can create false certainty: every required field forces an answer. If “missing,” “unknown,” “ambiguous,” or null is a legitimate state, model it explicitly. Otherwise the schema can turn uncertainty into fabricated precision.

This is closely related to the wider problem in AI hallucinations. Structure changes how an error is packaged; it does not remove the error. A wrong paragraph may invite skepticism. A wrong typed object may travel through several systems before anyone notices.

A Four-Layer Validation Model for Structured AI Responses

A dependable system validates structured AI output in layers. Each layer answers a different question. Combining them is more useful than searching for a single “reliability score.”

Paper-cut illustration of structured data passing through syntax, schema, business-rule, and evidence checkpoints before human review

A schema gate is one checkpoint in a broader validation line, not the finish line.

Layer 1: syntax and transport

Can the response be parsed? Did the request finish, or was it truncated, refused, timed out, or interrupted? Is the content type what the client expected? Did streaming finish cleanly? Even native structured modes need explicit handling for refusals and incomplete responses rather than pretending every request returns a normal object.

Layer 2: schema and type validation

Does the object match the exact schema version your application expects? Re-validate locally with the same source-of-truth type definition used by the application. Reject extra properties when they would hide drift. Treat schema versions like API versions, because renaming a field or changing an enum can break consumers even when the model behaves correctly.

Layer 3: domain and business rules

Now test rules the schema cannot fully express. A start date must precede an end date. Line items should reconcile to a total within a defined tolerance. A customer ID must exist. A selected plan must be available in the customer’s region. A medical code must be compatible with the source note and should never be auto-applied solely because a model returned it.

Layer 4: evidence, authority, and consequence

Verify claims against authoritative data. Require citations or source spans when extraction traceability matters. Confirm that the requesting user and acting service are authorized for the proposed operation. Decide whether the consequence is low enough to automate, or whether the result should become a review task. This is where structured data becomes an accountable decision rather than a convenient guess.

FailureBest detectorRecommended responseWhy retry alone is insufficient
Malformed or incomplete JSONParser and completion statusRetry with bounded attempts or use native schema modeA retry can fix transport or formatting, but not meaning
Schema mismatchLocal schema validatorReject, log, and retry with the validation error if appropriateRepeated failure may indicate unsupported complexity or a bad schema
Impossible valueDeterministic business ruleReject or route to reviewThe model may repeat a plausible but impossible value
Unsupported claimSource-span check or authoritative lookupMark unknown, request evidence, or escalateFormatting feedback does not supply missing evidence
Unauthorized actionExternal policy and identity layerBlock or request scoped approvalThe acting model must not grant itself authority
Low-confidence high-impact caseRisk policy plus human reviewQueue an evidence-rich reviewA second model answer is not independent authorization

For production systems, log which layer rejected the output. A single “model failed” metric hides the information needed to improve the prompt, schema, data source, validator, or workflow. The failure taxonomy should distinguish parse failure, schema failure, domain-rule failure, evidence failure, policy denial, and human rejection.

Where AI Structured Outputs Are Most Useful

The best use cases have a clear source, a bounded task, an explicit downstream consumer, and a validation path. Structured output is especially valuable when the alternative is brittle parsing of free-form prose.

Document extraction

Extract names, dates, clauses, totals, product codes, or obligations from documents. Include source spans or page references when auditability matters. Keep “not present” distinct from “inferred.” For financial or legal workflows, reconcile extracted values against authoritative records and route ambiguous documents to review.

Classification and routing

Classify support tickets, feedback, incidents, or content into a small set of operational categories. Enums are helpful because they prevent new labels from appearing silently. Add an unknown or needs_review path so difficult cases do not become forced guesses.

UI generation and form filling

A structured response can populate cards, timelines, comparison rows, or form fields. Keep rendering logic deterministic: the model supplies content and choices within an allowed schema; the application controls executable code, layout boundaries, sanitization, and accessibility.

Workflow state and agent handoffs

One agent or step can return a typed task state for the next step: objective, completed work, open questions, evidence references, risk flags, and required approvals. This reduces ambiguity between components. Pair it with the context practices in context engineering for AI agents, because a perfect handoff schema cannot compensate for missing or unauthorized context.

Evaluation records

Ask an evaluator to return a rubric outcome with separate fields for criteria, evidence, pass/fail status, and review notes. Do not treat model-generated scores as ground truth. Calibrate them against human-labeled examples and measure disagreement. The AI hallucination evaluation checklist shows how to test answers before users trust them.

When not to use a rigid schema

Free-form prose remains better for creative ideation, nuanced explanation, exploratory analysis, and situations where you do not yet understand the dimensions of the answer. A premature schema can narrow the model into categories that reflect your assumptions rather than the problem. Start with exploratory samples, discover the stable fields, and introduce structure only when a consumer truly needs it.

How to Design a Good JSON Schema for an AI Model

A good schema is small, explicit, versioned, and aligned with a real decision. It gives uncertainty somewhere honest to go. It does not ask the model to invent operational facts merely to satisfy required fields.

1. Begin with the consumer

List exactly what the next system needs. If a ticket router uses only category, severity, and review status, do not ask the model for fifteen decorative fields. Every extra field consumes attention, adds tokens, creates another place for error, and increases the chance that future code begins depending on an unreliable value.

2. Separate extraction from interpretation

Keep direct evidence apart from model judgment. For example, use quoted_deadline_text for the words found in the document, normalized_deadline for the interpreted date, and normalization_status for success, ambiguity, or missing context. This makes it possible to review the transformation rather than receiving only a polished conclusion.

3. Model uncertainty deliberately

Use nullable fields, explicit unknown states, ambiguity flags, and arrays that may be empty. Avoid asking for a confidence number unless you have a calibration method and a policy that uses it responsibly. A categorical reason such as “missing source evidence” can be more actionable than an invented decimal.

4. Prefer enums for operational decisions

Enums stop label drift. Define them narrowly and document each option. Include a fallback category. Do not put two different decisions into one enum—for example, mixing content category with escalation status. Separate fields are easier to validate and analyze.

5. Add descriptions that constrain meaning

Describe evidence requirements, units, reference times, and forbidden inference. “Total amount” is weaker than “Total payable amount printed on the invoice, in the currency shown; return null when not visible.” Descriptions should guide the model, while code enforces hard rules.

6. Keep nesting shallow

Deeply nested schemas are harder for models, developers, logs, and review tools. If a result becomes a miniature database, split the task into stages. Extract the core record first, validate it, then derive secondary structures in deterministic code or a second bounded model call.

7. Version the contract

Add an application-level schema version or bind the validator to a deployed version. Test old and new consumers during migrations. Save representative fixtures and adversarial examples. A prompt change, model change, schema change, and upstream document change can all alter outcomes even if the endpoint stays the same.

8. Refuse unsupported work cleanly

Design an explicit result for refusal, missing input, or insufficient evidence rather than forcing the normal success object. Your application should know the difference between “the answer is empty” and “the model could not safely answer.”

Editorial recommendation: if your schema cannot represent uncertainty, absence, refusal, and review, it is probably optimized for a demo rather than a dependable workflow.

Provider Support and Portability

OpenAI, Google, and Amazon Bedrock all document schema-constrained structured output, but their request shapes, supported models, supported JSON Schema features, streaming behavior, refusals, and tool combinations differ. Treat the full JSON Schema specification as a vocabulary, not proof that every keyword works everywhere.

For portability, keep one provider-neutral domain model in your application and write thin adapters for each model API. Generate provider schemas from the same source types when possible. Maintain a compatibility test suite that exercises required fields, enums, nulls, arrays, nesting, refusals, incomplete outputs, and deliberately ambiguous inputs.

Do not silently weaken validation when a provider lacks a schema feature. Move the constraint into application code and document the shift. For example, if numeric bounds are unsupported, still accept a numeric field at generation time but enforce the range after parsing. The same rule applies to string length, cross-field relationships, uniqueness, and external references.

Also test behavior rather than relying only on capability tables. A provider may guarantee schema adherence while different models vary in how well they choose correct values, respect field descriptions, use null, or avoid hallucinated evidence. Structural conformance is binary; task quality is empirical.

Interactive Decision Helper: Which Output Mode Fits?

Use this lightweight helper to choose a starting pattern. It does not replace provider documentation or a risk review; it clarifies whether your main need is human-readable prose, parseable data, an exact response contract, or a controlled action request.

Choose your workflow characteristics, then request a recommendation.

A common architecture uses more than one mode in the same turn. The model may request a tool through function calling, the application may execute a read-only lookup, and the model may then return a schema-constrained final answer for a UI. Keep each boundary explicit so logs show what was requested, what ran, what data came back, and what final result was accepted.

A Production-Minded Implementation Blueprint

You do not need an elaborate platform to start. You need a small contract, a representative test set, deterministic validation, and a clear fallback. Add complexity when observed failures justify it.

Implementation companion: Use the LLM structured output validation production checklist to turn this blueprint into concrete completion, schema, business-rule, evidence, security, retry, and release gates.
  1. Write the decision statement. Define what the structured result will influence and what it must never do automatically.
  2. Collect representative inputs. Include clean cases, incomplete cases, conflicting evidence, unusual formatting, adversarial text, and genuinely unanswerable examples.
  3. Design the smallest schema. Include unknown and review states. Separate extracted evidence from interpretation.
  4. Create a typed local model. Use a validation library in your application language and generate the API schema from that source when practical.
  5. Call a supported structured mode. Handle completion, refusal, truncation, timeout, and provider errors explicitly.
  6. Re-validate locally. Treat the provider response as untrusted input at the network boundary.
  7. Apply deterministic checks. Enforce cross-field rules, units, date logic, identifiers, totals, uniqueness, and reference-data lookups.
  8. Verify evidence. Require source spans or citations for claims that must be traceable. Confirm that cited text actually supports each value.
  9. Route by consequence. Auto-accept low-impact, well-validated cases; queue ambiguous or high-impact cases with the evidence a reviewer needs.
  10. Measure outcomes. Track parse success, schema success, domain-rule failure, evidence failure, reviewer disagreement, false acceptance, latency, cost, and drift by model and schema version.

A simple control-flow pattern

result = model.generate(task, source, schema)

if result.refused or result.incomplete:
    return route_to_fallback(result.status)

typed = schema_validator.parse(result.output)
domain_errors = check_business_rules(typed)
evidence_errors = verify_against_source(typed, source)

if domain_errors or evidence_errors:
    return review_queue(typed, domain_errors, evidence_errors)

if consequence_is_high(typed):
    return approval_queue(typed)

return accept(typed)

Notice what is missing: the model does not approve its own output, grant its own permissions, or decide that validation can be skipped. It proposes. Deterministic systems and accountable people decide what happens next.

Retries, repair, and fallbacks

A retry is appropriate for a transient failure, incomplete generation, or fixable validation error. Make retries bounded, record why they happened, and avoid feeding an endless loop with the same impossible task. If the source lacks a required fact, no number of retries can recover it. The correct result is missing information or human review.

Fallbacks can include a simpler schema, a different model, deterministic extraction, asking the user for missing information, or returning a safe partial result. A fallback should reduce uncertainty or impact, not merely produce something that passes the parser.

Testing that matters

Measure more than schema-pass rate. Build a golden dataset with expected values and acceptable alternatives. Include negative cases in which the correct behavior is null, unknown, refusal, or escalation. Track field-level precision and recall for extraction, classification confusion, evidence support, and reviewer disagreement. If a model returns 100 percent valid JSON and 12 percent wrong values, your parser dashboard is green while your product is failing.

Use the deployment practices in the AI agent deployment pipeline: version changes, run evaluations, release gradually, trace failures, and keep a rollback path. Even a non-agentic structured-output service benefits from the same discipline.

Common Misconceptions

“The API guarantees 100 percent reliability.”

A provider may guarantee adherence to a supported schema for a completed, non-refused response. That is structural reliability. It is not a guarantee of correct extraction, correct classification, truthful content, or successful business outcomes.

“If the output validates, we can store it as truth.”

Validation means the object fits rules you defined. Store provenance, model and schema versions, source references, validation status, and reviewer outcome when the record may later be audited. Separate model-derived data from authoritative system-of-record fields.

“A confidence field tells us when the model is right.”

Not automatically. A requested confidence number may be another generated value. Use empirical calibration against labeled outcomes before assigning thresholds. Even calibrated confidence does not replace authorization or impact-based review.

“Function calling means the model executes the function.”

In the common client-side pattern, the model proposes a function name and arguments. Your application decides whether and how to run it, then returns the result. Some platforms offer server-side tools, but the security principle remains: execution and policy must be controlled by the surrounding system.

“A bigger schema captures more value.”

A bigger schema often creates more failure surfaces. Include fields because a consumer needs them, not because the model can fill them. Derive deterministic fields in code and split large tasks into testable stages.

Final Takeaway: Structure the Response, Validate the Meaning

AI structured outputs are one of the most useful bridges between language models and dependable software. They replace prompt-only formatting tricks with explicit contracts. They make responses easier to parse, test, render, route, compare, and monitor. They also make errors look more like ordinary data, which is why the reliability boundary must be understood clearly.

Use free-form text when a person needs explanation. Use JSON mode when parseable JSON is enough and exact fields can remain flexible. Use schema-constrained output when an application needs a stable response contract. Use function calling when the model must ask your system to retrieve information or perform an action. In every case, validate according to consequence.

The mature pattern is not “model returns JSON, therefore the workflow is safe.” It is: model returns a typed proposal; the application validates structure; deterministic code checks domain rules; authoritative sources verify evidence; policy checks authority; and a person reviews the cases whose impact or ambiguity demands judgment.

Next step: choose one narrow extraction or classification task, design a schema that represents uncertainty honestly, build twenty difficult test cases, and measure semantic correctness—not just JSON validity. Then explore the related AI CORE guides above before connecting the result to automated actions.

FAQ About AI Structured Outputs

What are AI structured outputs in simple terms?

They are AI responses forced into a predictable machine-readable shape, usually defined with JSON Schema. The schema can require specific fields, types, and allowed values so an application can parse the response reliably.

What is the difference between JSON mode and structured outputs?

JSON mode ensures the response is valid JSON. Structured outputs go further by requiring the JSON to match a defined schema. Neither one guarantees that the values are factually or semantically correct.

Are structured outputs the same as function calling?

No. Structured outputs usually shape the model’s final response. Function calling gives the model a structured way to request that your application use a tool or function. Both may use JSON Schema, but function calling can lead to an external action.

Can schema-valid AI output still hallucinate?

Yes. A model can produce an object that perfectly matches the schema while inventing an identifier, misreading a date, choosing the wrong category, or asserting a value unsupported by the source. Validate evidence and business rules separately.

Should I validate structured output in my application?

Yes. Re-validate the response locally, check domain rules and authoritative records, handle refusals and incomplete outputs, and route ambiguous or high-impact cases to review. Treat the response as untrusted network input.

How should a schema represent uncertainty?

Allow null or explicit values such as unknown, ambiguous, missing evidence, and needs review. Do not force a binary answer when the source may not support one. Keep extracted evidence separate from model interpretation.

When should I avoid structured outputs?

A rigid schema may be a poor fit for brainstorming, creative writing, exploratory analysis, or a new problem whose stable dimensions are not yet understood. Use free-form samples first, then add structure when a downstream consumer needs it.

Do all model providers support the same JSON Schema features?

No. Providers and models support different subsets and request formats. Maintain a provider-neutral domain type, adapter tests, and local validation. Check current official documentation before depending on a particular schema feature.

What metrics should I track in production?

Track completion and refusal rates, parse and schema success, field-level correctness, domain-rule failures, evidence failures, reviewer disagreement, false acceptance, latency, cost, and drift by model, prompt, and schema version.