Back to blog
Engineering·October 8, 2026·11 min read

Schema-Based AI JSON Extraction: How outputSchema Actually Behaves

How outputSchema really behaves across OpenAI, Anthropic, Gemini and local engines: what constrained decoding guarantees, where it stops, and how Anakin.io's URL Scraper and Wire turn schema-based extraction into a dependable contract for real-world web data.

A

Arun Singh

Anakin Team

outputSchema is the contract. Anakin handles the web behind it: Fetch, Clean and Extract run on Anakin, then Validate runs in your code.

Most teams start the same way. You prompt the model: "Return only valid JSON matching this shape." You add "Do not include markdown" and "Do not explain." It works in the demo. Then production arrives with missing required fields, extra keys, wrong types, nested objects collapsed into strings, and the classic "Sure, here's the JSON:" prefix.

Schema-based extraction (what people mean when they write outputSchema, response_format: { type: "json_schema" }, or pass a Pydantic or Zod model) moves this out of prompt engineering and into constraints. That's a genuine step forward. Getting the most out of it means knowing what each provider actually enforces, and what's still yours to handle.

What "outputSchema" actually is

In LLM APIs and the frameworks built on them (OpenAI, Anthropic, Gemini, Instructor, the Vercel AI SDK and others), outputSchema is a JSON Schema, or something that compiles to one, that the runtime hands to the model provider. The provider is responsible for making the output conform.

There are three mechanisms under the hood, and they come with very different guarantees.

1. Prompt + post-validation. The schema is rendered as text into the system prompt or a tool description. The model can still generate anything; you parse, validate and retry on failure. Many library fallbacks and local setups work this way. It's much better than pure prompting, but it's still probabilistic.

2. JSON mode (syntax-constrained). OpenAI's response_format: { type: "json_object" } and similar modes constrain the output to syntactically valid JSON. It will parse, but nothing guarantees that it matches your schema.

3. Constrained decoding (strict structured outputs). The schema is compiled into a grammar or state machine. At every generation step, the sampler masks any token that would make the output invalid. OpenAI's strict: true, Anthropic's structured outputs (output_config.format and strict tool use), Gemini's responseSchema and responseJsonSchema, and local engines such as vLLM and SGLang all work this way. On a normal completion, the output is guaranteed to parse and match the schema: no "almost JSON", no missing required keys, no invented enum values.

Most frameworks use the third path when the provider supports it and fall back to the first otherwise. That fallback is why the same code can feel rock-solid on one model and flaky on another.

A note on MCP. MCP tools also declare an outputSchema, but it plays a different role. It describes the structuredContent that a tool server returns from its own code, so clients can validate it. No model decoding is involved.
Diagram comparing three ways to get structured JSON from an LLM: prompt plus validation, JSON mode, and constrained decoding, the only method that guarantees schema-based JSON extraction
Token masking is the only mechanism that guarantees the schema.

The behavioral details that matter

The root is an object

OpenAI's strict mode and MCP tool output both require a top-level object, and it's the safe default everywhere. If your natural shape is a list, wrap it: { "items": [...] }.

Each provider speaks its own dialect of JSON Schema

If you pass a rich, general-purpose schema straight through, expect 400 errors or silently dropped constraints. The rules worth knowing:

  • additionalProperties: false is required, not implied. OpenAI's strict mode and Anthropic both expect it set explicitly on every object.
  • OpenAI's strict mode requires every property in required. Model optional fields as a union with null: "type": ["string", "null"]. This is the most common reason a Zod or Pydantic schema fails when you pass it straight through.
  • Support for recursion varies. OpenAI supports recursive schemas through $ref and $defs, and so does Gemini's responseJsonSchema on 2.5+ models. Anthropic doesn't.
  • Composition is limited. anyOf is the portable choice for unions, though OpenAI doesn't allow it at the root. Support for oneOf, allOf and not varies by provider.
  • Value constraints may be enforced after the fact. For example, Anthropic doesn't enforce minimum, maximum, multipleOf, minLength or maxLength during decoding. Its official SDKs move them into field descriptions and validate them on the client. Treat range and length checks as validation, not as generation guarantees.

The schema you hand the model is rarely identical to the full schema you'd use for ordinary validation, and that's fine. Keep both, generated from one source.

Field descriptions are your extraction instructions

OpenAI, Anthropic and Gemini pass the schema into the model's context, descriptions included. Once the structure is constrained, the descriptions are the best place for guidance: "price as a number, without the currency symbol" beats a paragraph of system prompt.

One caveat applies to local guided decoding (vLLM, SGLang, Ollama's format). There the schema typically drives only the token mask, and the model never sees your descriptions unless you also put the schema in the prompt. Ollama's docs recommend doing exactly that.

Property order is generation order

The model writes fields in the order the schema lists them. Put supporting fields (evidence, source text, reasoning) before the conclusions that depend on them. On Gemini's responseSchema, set propertyOrdering to control this.

The first request with a new schema is slower

The grammar has to be compiled before it's cached. Anthropic, for instance, caches compiled grammars for 24 hours, and editing only a name or description doesn't invalidate the cache, so you can iterate on descriptions freely. The one-off compile shows up as occasional latency spikes that look like the model "thinking harder."

Refusals and truncation are the two exits from the guarantee

  • Refusals. OpenAI returns them in a dedicated refusal field, and Anthropic returns stop_reason: "refusal". Either way, the output may not match your schema.
  • Truncation. Constrained decoding masks tokens, but it can't make the model finish early. If you hit the token limit, you get a cut-off JSON prefix that won't parse. Check finish_reason: "length" (OpenAI) or stop_reason: "max_tokens" (Anthropic) before parsing, and size the limit for your largest realistic output.

Schema adherence tells you the output is well-formed. It doesn't tell you the content is right.

Diagram showing constrained decoding guarantees schema-valid JSON output except for two exits, a model refusal with stop_reason refusal or token truncation with finish_reason length
Check finish_reason / stop_reason before you parse the response.

Provider cheat sheet

ProviderNative mechanismWhat to know
OpenAIresponse_format / text.format with type: "json_schema" and strict: true; strict function callingRoot must be an object. Every property is required (use null unions for optional fields), and additionalProperties: false is required. Recursion is supported via $ref/$defs. Refusals arrive in a dedicated refusal field.
Anthropicoutput_config.format with a JSON schema; strict: true on toolsConstrained decoding, generally available. No recursive schemas. Numeric and string-length constraints are validated client-side by the SDKs. Refusals arrive as stop_reason: "refusal". Compiled grammars are cached for 24h.
GeminiresponseSchema (OpenAPI subset) or responseJsonSchema (JSON Schema, 2.5+)Use responseJsonSchema for unions (anyOf) and recursion. Use propertyOrdering with responseSchema.
Local (vLLM, SGLang, Ollama)Grammar- or FSM-constrained decodingThe schema usually isn't shown to the model, so add it to the prompt. Quality depends on the engine and the model size.

Frameworks hide the wiring, but they don't erase these differences. Test the exact schema against the exact model you'll run in production.

Where libraries still earn their keep

Native constrained decoding gives you structural correctness. It doesn't give you:

  • business-rule validation ("end date must be after start date")
  • cross-field consistency
  • automatic retries that feed the validation error back to the model
  • one API across providers

That's why Instructor and similar libraries remain valuable. They use the native path where it exists, fall back to tool calling or JSON mode where it doesn't, and add the retry loop that turns "schema-valid but semantically wrong" into something recoverable.

The real-world problem: the schema is the easy part

Everything above assumes you already have clean text to extract from. On the web, you usually don't. Before a model ever sees your schema, something has to:

  • fetch a page that renders in JavaScript, sits behind bot protection, or changes by country
  • turn messy HTML into something a model can read reliably
  • keep working when the site ships a redesign
  • notice, for recurring jobs, when the data you care about actually changes

That's the part of the pipeline Anakin.io takes off your plate. It offers the schema-based contract in two forms: you write the schema (the URL Scraper), or Anakin already has (Wire).

Comparison of Anakin's URL Scraper, where you bring your own outputSchema for any URL, against Wire's prebuilt catalog of 953 sites with the schema already written
Long-tail pages, write it once. Popular sites, don't write it at all.

URL Scraper: bring your own schema

The URL Scraper fetches any page, with optional stealth headless-browser rendering and proxies in the country you choose, and then runs an AI extraction pass over the content. Call POST /v1/url-scraper and poll GET /v1/url-scraper/{id}, or use POST /v1/url-scraper/scrape to get the result inline. POST /v1/url-scraper/batch takes up to 10 URLs at once.

It supports two extraction modes:

  • Inferred structure. Set generateJson: true, or add "json" to formats. The completed job includes a generatedJson object with whatever fields the model found useful, such as title, price, author, tags or a summary. It's ideal for exploration and one-off pages where you don't know the fields yet.
  • Your schema. Pass an outputSchema: a JSON Schema that describes exactly the fields, types and nesting you want. It implies generateJson: true, and generatedJson comes back with the fields your schema describes. Downstream code gets a contract instead of a best-effort blob.
curl -X POST https://api.anakin.io/v1/url-scraper/scrape \
  -H "X-API-Key: $ANAKIN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://shop.example.com/products/123",
    "outputSchema": {
      "type": "object",
      "properties": {
        "title":    { "type": "string",           "description": "Product name from the main heading, without the brand prefix" },
        "price":    { "type": ["number", "null"], "description": "Current selling price as a number, no currency symbol; null if not shown" },
        "currency": { "type": ["string", "null"], "description": "ISO 4217 code, e.g. USD" },
        "inStock":  { "type": "boolean",          "description": "true only if the page lets you add the item to the cart right now" }
      },
      "required": ["title", "price", "currency", "inStock"],
      "additionalProperties": false
    }
  }'

A few habits make this rock-solid:

  • Author the schema once, in code. Define it in Zod or Pydantic, export it to JSON Schema, and follow the portable rules above: closed objects, explicit nullables, shallow nesting. That keeps the schema compatible with every provider's dialect.
  • Validate at your boundary with the same model. One definition serves as both the request contract and the response check, which makes the pipeline type-safe end to end. It's good practice with any extraction step.
  • Let descriptions do the instructing. They're the most direct way to tell the extractor what you mean by "price" or "in stock."
  • There are no selectors to break. The schema describes the data, not the markup, so a redesigned page doesn't mean rewriting CSS paths.

Monitoring: watch fields, not pages

Anakin monitors can diff a whole page, or, with watchMode: "specific_data" and an outputSchema, track only the fields you care about, such as price, stock status or a filing date. A layout tweak doesn't change the extracted fields, so it doesn't trigger an alert, while a price drop does. Turn on aiMode to filter out trivial noise and get a summary of what actually changed.

Diagram of Anakin Monitoring watching only schema-defined fields like price and stock_status with watchMode specific_data, instead of diffing the entire page with watchMode full_page
A redesign changes the layout, not price or stock_status.

Wire: the schema is already written

Wire is a catalog of pre-built actions for well-known sites. As of October 2026 it covers 957 sites and more than 5,000 actions, from Amazon, Walmart and Zillow to Beatport, GitHub, and public-data sources such as UK Companies House and ClinicalTrials.gov.

Each action is a packaged, versioned extractor for one specific job: "Beatport → Track Details", for example, or Amazon product search. Using one takes two steps:

  1. Discover. GET /v1/wire/catalog lists the sites. GET /v1/wire/catalog/{slug} returns each action's parameter schema, whether it reads or writes, its auth mode and its credit cost. GET /v1/wire/resolve (alias /search) finds the right action by intent.
  2. Run. POST /v1/wire/task returns a job ID; poll GET /v1/wire/jobs/{id} for the structured result.

There's no extraction prompt and no schema to write. The action defines the output shape, and Anakin maintains the action when the site changes, so your integration keeps receiving the same structure. If a site you need isn't covered, POST /v1/wire/build-request asks Anakin to build an action for it.

Wire also plugs into monitoring. A Wire-scoped monitor runs an action on a schedule and can diff just the JSON paths you choose.

Who authors the contract?

Both paths rest on the same idea: schema-based extraction turns "please give me JSON" into a contract. The difference is who writes it.

URL Scraper + outputSchemaWire action
Who defines the shapeYouAnakin
Works onAny URLCatalog sites (or request one)
Best forLong-tail pages, custom fieldsPopular sites, recurring pulls
When the site changesThe schema describes data, not markupAnakin updates the action
Recurring jobsspecific_data monitorWire-scoped monitor

Practical takeaways

  • Use the native constrained path when your schema fits the provider's dialect. It's the only real guarantee.
  • Keep schemas closed, explicit and shallow: set additionalProperties: false, make nullability explicit, and check recursion and composition support per provider before relying on them.
  • Put extraction guidance in field descriptions. On local engines, put the schema in the prompt too.
  • Order fields so supporting evidence comes before conclusions.
  • Handle refusals and truncation explicitly. A schema-valid response isn't necessarily a successful extraction.
  • Test the exact schema against the exact model and provider you'll run in production.
  • Reach for a library such as Instructor when you need business-rule validation and retries.
  • For web data, let the infrastructure do the heavy lifting. Use Wire when the site is in the catalog, outputSchema on the URL Scraper when it isn't, and field-level monitors for anything recurring. Validate at your boundary with the same model you authored the schema from.

Schema-based extraction removed the most annoying class of parsing bugs. The outputSchema parameter is the contract. The guarantees come from what enforces it: token masking in the model, validation at your boundary and, for web data, an extraction layer like Anakin's that keeps delivering the same shape when the page underneath it changes.

If you're tired of maintaining selectors and parsing almost-JSON, start with Anakin: point the URL Scraper at a page with your outputSchema, or pick a Wire action from the catalog, and see structured output before you write a line of parsing code.

Frequently Asked Questions

What is outputSchema in an LLM API call?

outputSchema is a JSON Schema, or something that compiles to one, that you hand to a model provider so it constrains its output to match. Depending on the provider, this is enforced either through constrained decoding, where the sampler masks any token that would break the schema during generation, or through weaker prompt- or syntax-based methods that only guarantee parseable JSON, not a schema match. OpenAI, Anthropic and Gemini all support the strict, constrained-decoding version today.

Does a schema-valid response mean the extracted data is correct?

No. Constrained decoding guarantees the output parses and matches your schema's structure and types. It does not guarantee the values are factually correct, that business rules like "end date must be after start date" hold, or that the model picked the right field from the page. Schema adherence and extraction accuracy are two separate problems, and only the first one is solved at the decoding level.

Which LLM providers support strict structured output enforcement?

OpenAI (strict: true), Anthropic (output_config.format and strict tool use) and Gemini (responseSchema and responseJsonSchema on 2.5+ models) all run genuine constrained decoding today. Local inference engines such as vLLM and SGLang support it too, through backends like xgrammar, outlines and llguidance. Each provider enforces a different subset of JSON Schema, so a schema written for one needs portable rules, a closed object, explicit nullables, shallow nesting, to work across all of them.

How does Anakin handle schema-based extraction from live web pages?

Two ways, depending on whether the fields you need already exist as a catalog action. With the URL Scraper, you pass your own outputSchema against any URL and get generatedJson back shaped to it. With Wire, the schema is already written for you as a pre-built, versioned action across the site catalog, so there's no schema to author and no selector to maintain when the site changes.

Can you monitor a webpage for changes to specific fields instead of the whole page?

Yes. Anakin's Website Monitoring supports watchMode: "specific_data" with an outputSchema, so the monitor diffs only the fields you define, like price or stock status, instead of the whole rendered page. A layout or copy change on the page does not fire an alert; a change to one of the fields you actually schema-defined does. Turning on aiMode additionally filters trivial changes out of the alert summary.