This lesson on Prompting + Structured Extraction is hands-on and example-driven. You will learn how to reliably instruct a Large Language Model (LLM) to return data in a specific, machine-readable JSON format. You will also master the discipline of grounding, ensuring the LLM refuses to answer when the prompt falls outside its defined scope or schema. This skill is crucial for building robust, predictable RAG agents.
What You'll Be Able To Do
- Define a Pydantic schema for required data extraction fields.
- Write a function that uses an LLM API to enforce JSON output based on a provided schema.
- Implement a system prompt that mandates clean refusal for out-of-distribution queries.
- Validate extracted JSON against the expected schema programmatically.
- Apply grounding discipline to limit LLM responses strictly to provided context or defined tasks.
Detailed Concept Walkthrough
1. Structured Extraction via JSON
LLMs must return data in structured formats like JSON for reliable downstream processing in agent workflows. Using a defined schema ensures the output is predictable and machine-readable.
- Mechanism: Modern LLM APIs offer a 'JSON mode' or 'response_format' parameter, which constrains the model's output token generation to adhere to valid JSON syntax. This constraint is applied at the token level during the decoding phase.
- Best Practice: Always provide the model with an explicit JSON schema (e.g., Pydantic schema converted to JSON schema) within the system prompt to define required fields, types, and descriptions. This minimizes ambiguity and improves extraction accuracy.
- Under the Hood: The model's decoder layer uses grammar constraints (like JSON Schema) during inference, preventing the generation of tokens that would break the structure. This significantly improves reliability over simple text prompting, which often results in malformed JSON due to conversational drift.
from pydantic import BaseModel, Field
import json
# 1. Define the required structure
class UserProfile(BaseModel):
name: str = Field(description="Full name of the user.")
age: int = Field(description="User's age in years.")
# 2. Convert schema for the API
schema_json = UserProfile.model_json_schema()
# Conceptual API call setup:
# response = llm_client.generate(
# prompt="Extract data for John Doe, 35.",
# response_format={"type": "json_object", "schema": schema_json}
# )
Key Takeaway: Use explicit JSON mode and provide a detailed schema to guarantee predictable, machine-readable output from the LLM.
2. Grounding Discipline
Grounding is the principle of limiting the LLM's response space, forcing it to rely only on provided context or strictly defined instructions. This prevents hallucination and scope creep, ensuring agent reliability.
- Best Practice: The system prompt must clearly define the boundaries of the task, explicitly stating what information sources are valid and what actions are permissible. For RAG, this means instructing the model to only use the retrieved documents.
- Execution Flow: When processing a query, the LLM first checks if the necessary information is present in the provided context or if the query aligns with the extraction schema before attempting to generate an answer. If the information is absent, the model must follow the refusal instruction.
- Mechanism: In structured extraction, grounding means the model must only populate the JSON fields with data found in the input text, or return a specific refusal object if the data is missing. This prevents the model from "hallucinating" data to fill empty fields.
Key Takeaway: Strict grounding ensures the LLM acts as a reliable data processor within defined constraints, not a general knowledge engine.
3. Out-of-Distribution Refusal
OOD refusal is the ability of the LLM to cleanly reject prompts that fall outside the scope of its current task or schema. This structured rejection mechanism builds essential robustness into the agent.
- Syntax Rule: The system prompt must include a mandatory refusal instruction, often requiring a specific refusal phrase or a predefined JSON structure (e.g.,
{"status": "refused", "reason": "..."}). This instruction must be prioritized over general conversational abilities. - Under the Hood: The model prioritizes the refusal instruction when the input prompt conflicts with the extraction task (e.g., asking a general knowledge question when the task is strictly data extraction). This is a direct result of fine-tuning models for instruction following.
- Best Practice: Test OOD refusal using a "golden suite" of known out-of-scope questions (like the 20Q suite) to ensure the model consistently returns the expected refusal format rather than attempting a creative answer. Consistent refusal is a key metric for agent robustness.
from pydantic import BaseModel, Field
# Define the mandatory refusal structure
class Refusal(BaseModel):
status: str = Field(default="REFUSED")
reason: str = Field(description="Why the query is OOD.")
# System Prompt Excerpt:
# "If the input text does not contain the required data,
# you MUST return a JSON object matching the Refusal schema
# instead of the UserProfile schema."
Key Takeaway: Mandating a structured refusal mechanism is essential for handling unexpected inputs gracefully and maintaining agent stability.
Topics Covered in Prompting + Structured Extraction
- Intro to Structured Output (0:00 - 1:30) — Structured output is necessary for agents to reliably consume LLM responses.
- Defining the Schema (1:30 - 3:00) — Pydantic models are used to define the required data structure and generate the necessary JSON schema.
- Enforcing JSON Mode (3:00 - 4:45) — The LLM API must be configured with JSON mode and the schema to guarantee valid output.
- Grounding and Refusal (4:45 - 6:30) — Grounding discipline requires the model to strictly adhere to the task and refuse OOD queries.
- Testing OOD Robustness (6:30 - 8:00) — Testing with out-of-scope questions ensures the structured refusal mechanism functions correctly.
GenAI + RAG Agents (DS Lite) Cheat Sheet
-
response_format={"type": "json_object"}— Enforces valid JSON output from the LLMllm.generate(..., response_format=...) -
Pydantic BaseModel— Defines required data structure and typesclass Data(BaseModel): field: str -
model_json_schema()— Converts Pydantic model to API-ready schemaschema = Data.model_json_schema() -
System Prompt Refusal— Mandates specific output for OOD queries"If OOD, return {'status': 'REFUSED'}" -
Grounding Discipline— Limits LLM to provided context/task- In practice: Strictly use the provided text for extraction.
Comparison Table
| Aspect | Standard Text Prompting | Structured JSON Prompting |
|---|---|---|
| Goal | General response generation | Specific data extraction |
| Output Format | Unpredictable natural language | Validated JSON object |
| Downstream Use | Requires complex parsing | Direct machine consumption |
| Robustness | High risk of hallucination | Enforces schema and refusal |
Common Pitfalls
- Mistake: Relying on text instructions alone to produce JSON.
Avoid: Always enable the API's dedicated
response_formator JSON mode parameter. - Mistake: Not defining a specific refusal mechanism for OOD queries. Avoid: Include a mandatory, structured refusal instruction in the system prompt.
- Mistake: Using vague field descriptions in the Pydantic schema.
Avoid: Use
Field(description=...)to clearly define the expected content and format for every field. - Mistake: Allowing the LLM to invent data if it's missing (poor grounding). Avoid: Explicitly instruct the model to only use source material or return a refusal object.
FAQs
- Why is using JSON output better than just asking for a list in text? JSON output guarantees machine readability and type consistency, allowing direct parsing into Python objects without fragile regex or NLP parsing steps.
- What is an Out-of-Distribution (OOD) prompt? An OOD prompt is a query that falls outside the scope of the task defined in the system prompt, such as asking for general knowledge during a strict data extraction task.
- How does Pydantic help with structured extraction? Pydantic defines the required data types and structure (schema) upfront, and its schema can be passed directly to the LLM API to enforce output validation and consistency.
- What is the "grounding pattern" mentioned in the source notes? The grounding pattern refers to the disciplined approach of limiting the LLM's knowledge base strictly to the provided context or task definition, often enforced via system prompts.