Structured Output That Holds in Production
"Just ask for JSON" is why your pager went off. Structured output is where the model's unpredictable text meets your typed code. Get this boundary wrong and every downstream system inherits the model's failure modes.
Here's the pattern that works and the traps we've watched teams fall into.
The three ways to get structured output, ranked
1. Native structured output APIs (best). OpenAI's response_format with a JSON schema. Anthropic's tool use with strict: true. Google's response schemas. When available, use these. They constrain generation at the token level, so the model emits valid structure.
2. Function calling with a schema (second best). Same idea without token-level enforcement, so outputs can still deviate from the schema. The tool call is your structured output; you don't execute a tool. This still works reliably and is universal across providers.
3. Prompted JSON with post-hoc parsing (last resort). "Return your answer as JSON matching this schema." Model tries. Sometimes succeeds. Fails in interesting ways under load, with long context, or when the schema is complex. Use only when the other two aren't available.
The pattern that holds up
import { z } from "zod"; const OrderClassification = z.object({ category: z.enum(["refund", "shipping", "product", "other"]), priority: z.enum(["low", "medium", "high", "urgent"]), needs_human: z.boolean(), reasoning: z.string().max(280),}); type OrderClassification = z.infer<typeof OrderClassification>; async function classifyOrder(text: string): Promise<OrderClassification> { const response = await anthropic.messages.create({ model: MODEL, // the snapshot you validated against tools: [{ name: "classify_order", description: "Classify a customer support message about an order.", input_schema: zodToJsonSchema(OrderClassification), strict: true // token-level schema enforcement (option one) }], tool_choice: { type: "tool", name: "classify_order" }, messages: [{ role: "user", content: text }] }); const toolUse = response.content.find(c => c.type === "tool_use"); if (!toolUse) throw new StructuredOutputError("no tool call"); return OrderClassification.parse(toolUse.input);}Four things this pattern does right. It forces the model to invoke a tool (no free-form response) and asks the provider to enforce the schema at the token level; drop strict on a provider that lacks it and you have option two, with the same code and the same parse call. It uses a schema library (Zod, but Pydantic or Effect Schema work the same way). It parses and validates on the way out, so downstream code sees a fully typed object. And it throws a specific error type so the caller can decide how to handle failures.
The traps
Optional fields the model fills with garbage. If a field is optional in your schema, some models fill it with a plausible value rather than omitting it. Make fields non-optional and use enums with an explicit unknown or not_applicable value; Tool Design Is Prompt Design makes the same case for tool arguments.
Long enums degrade accuracy. An enum with 40 categories will get miscategorized. If you need 40 categories, hierarchically classify: first choose one of 5 broad categories, then classify within that. Two smaller decisions beat one big one.
Nested objects fail more than flat ones. Deeply nested schemas are harder for the model to fill correctly. Flatten when you can. If you have three related fields, make them three top-level fields.
Free-text fields are where injection sneaks through. A reasoning field is useful for eval and debugging, but it's also a place where prompt injection can leak through. Treat any string field as untrusted input to the next system.
Validation is the boundary
The Zod parse call above is the contract enforcement point between the model and your typed code. What to do when validation fails matters more than the schema itself.
Three strategies, ranked by production stability:
Retry with schema in the error. Send the validation error back as a user turn or tool result and ask for a fix. It works well. Cap at one retry, then fail.
Fall back to a smaller schema. If the full schema fails, try a minimal version with only required fields. Return partial data with a flag. Useful when downstream can handle partial data.
Fail fast to a human queue. For anything high-stakes, don't retry. Route the message to a human review queue with the raw model output attached. This is the right answer for regulated environments.
Never coerce a failure into a value. Never ?? null your way past a validation failure. The model gave you something invalid; downstream systems should know.
What to log
Every validation failure. Every retry. The raw model output alongside the parsed result. The prompt and tool schema versions. When your fraud model or your ticket router starts behaving oddly, this trail is what tells you the structured output has been degrading for two weeks.
Related
Working on something like this?
Start a Conversation