This lesson on FastAPI Serving (/health, /predict, Schemas) is hands-on and example-driven. You will learn how to build production-ready ML inference APIs using FastAPI. You will implement robust data validation using Pydantic schemas for both request payloads and responses, ensuring your /predict endpoint is reliable and handles bad input gracefully.
What You'll Be Able To Do
- Define Pydantic models for structured request and response payloads.
- Implement a /health endpoint that reports service status and version information.
- Write a FastAPI route that accepts POST requests and performs input validation.
- Handle and interpret automatic 422 Unprocessable Entity responses from FastAPI.
- Integrate model inference logic into the validated request flow.
- Configure response modeling to guarantee output structure for clients.
Detailed Concept Walkthrough
1. Pydantic Schemas: Defining Data Contracts
Pydantic defines the expected structure and types for data entering and leaving the API, acting as a strict contract between the client and server. This ensures data integrity and provides automatic documentation and validation.
- Mechanism: You define schemas by inheriting from
pydantic.BaseModeland specifying fields using standard Python type hints. Pydantic validates incoming JSON against these types and constraints upon request. - Under the Hood: Validation happens immediately upon receiving the request body. If validation fails, Pydantic raises a
ValidationError, which FastAPI catches and converts into a standardized 422 response before your route logic executes. - Best Practice / Nuance: Define separate schemas for input (
InputData) and output (PredictionResult) to manage data transformation, enforce constraints specific to the model, and prevent exposing internal model details to the client.
from pydantic import BaseModel, Field
# Define the expected input structure and constraints
class PredictionInput(BaseModel):
feature_a: float = Field(..., description="First input feature, must be positive", gt=0)
feature_b: int = Field(1, description="Second input feature, defaults to 1")
# Define the expected output structure
class PredictionResult(BaseModel):
prediction_id: str
result: float
Key Takeaway: Pydantic schemas enforce type safety and structure, eliminating boilerplate validation code and standardizing API contracts.
2. Service Health and Versioning
The /health endpoint is a standard operational check used by load balancers and monitoring systems to confirm the service is running and responsive. It must be lightweight and return status information quickly.
- Mechanism: Implement a simple
GETroute (@app.get("/health")) that returns a dictionary or a Pydantic model containing status and version details. This confirms the application is initialized and accepting requests. - Execution Flow: This endpoint should bypass heavy processing, database lookups, or model inference. Its sole purpose is to confirm the FastAPI application, Uvicorn server, and underlying Python environment are operational.
- Best Practice / Nuance: Include the service version (e.g.,
"v1.0.0") in the health response. This version should be sourced from a configuration file or environment variable, ensuring consistency across deployed containers.
from fastapi import FastAPI
app = FastAPI()
SERVICE_VERSION = "1.0.0" # Loaded from config/env in production
@app.get("/health", tags=["Monitoring"])
async def health_check():
# Returns 200 OK by default
return {
"status": "ok",
"version": SERVICE_VERSION
}
Key Takeaway: A fast, reliable /health check is mandatory for cloud deployment readiness, automated scaling, and zero-downtime deployments.
3. Validated Prediction Routes
The /predict endpoint handles the core ML inference logic, accepting structured input data via a POST request and returning predictions. FastAPI automatically handles the deserialization and validation of the request body using the specified Pydantic model.
- Mechanism: Define the route using
@app.post("/predict"). The function signature takes the Pydantic input model as an argument (e.g.,data: PredictionInput). FastAPI uses this type hint to validate the incoming JSON payload. - Execution Flow: If the incoming JSON matches the
PredictionInputschema, FastAPI deserializes it into a Python object (data) and passes it to the function. If validation fails, the function is never called, and a 422 response is sent. - Best Practice / Nuance: Use the
response_modelargument in the decorator (@app.post(..., response_model=PredictionResult)) to ensure the output data structure conforms to the defined contract, even if your internal logic returns extra fields.
from typing import List
# Assume PredictionInput and PredictionResult are defined
# Placeholder for model loading
MODEL = object()
@app.post("/predict", response_model=PredictionResult)
async def predict_route(data: PredictionInput):
# 1. Input is guaranteed to be valid here
input_features = [data.feature_a, data.feature_b]
# 2. Perform inference (simplified)
model_output = 0.5 * data.feature_a + data.feature_b
# 3. Return the result conforming to PredictionResult schema
return {
"prediction_id": "abc-123",
"result": model_output
}
Key Takeaway: Type-hinting the route parameter with a Pydantic model enables automatic request body validation and deserialization, simplifying the route logic.
4. Automatic 422 Error Handling
When a client sends data that violates the defined Pydantic schema (e.g., wrong type, missing required field), FastAPI automatically returns a 422 HTTP status code (Unprocessable Entity). This standardizes error reporting for input failures.
- Mechanism: Pydantic's validation failure generates a detailed JSON error response body that adheres to the standard OpenAPI format. This body specifies exactly which fields failed validation and provides context (e.g., 'value is not a valid float').
- Under the Hood: FastAPI's exception handlers intercept the
ValidationErrorraised during the request parsing phase. It then constructs the standardized 422 response, preventing the malformed data from ever reaching your application code. - Best Practice / Nuance: Clients should be programmed to specifically look for and parse the 422 response body. This allows them to correct the input data and retry the request, rather than treating it as a generic server failure (like a 500).
{
"detail": [
{
"loc": [
"body",
"feature_a"
],
"msg": "value is not a valid float",
"type": "type_error.float"
}
]
}
Key Takeaway: The 422 status code signals a client-side data error, providing structured, machine-readable feedback for input correction.
Topics Covered in FastAPI Serving (/health, /predict, Schemas)
- Lesson Objectives (0:00 - 1:00) — Introduction to building a robust, validated ML serving API using FastAPI and Pydantic.
- Pydantic Schema Definition (1:00 - 3:00) — Defining input and output data structures using BaseModel and type hints to establish the API contract.
- /health Endpoint Setup (3:00 - 4:30) — Implementing the basic GET route for service health checks, including version reporting.
- /predict Route Implementation (4:30 - 7:00) — Setting up the POST route and using Pydantic models as type-hinted parameters for automatic request validation.
- Testing 422 Validation (7:00 - 9:00) — Demonstrating how FastAPI automatically rejects malformed input and returns a structured 422 error response.
- Response Modeling (9:00 - 11:00) — Applying the response_model argument to guarantee the output structure and prevent data leakage.
- Summary and Next Steps (11:00 - 12:00) — Reviewing the core components and preparing the API for integration with smoke tests in the next lesson.
MLOps + Cloud Deploy (AWS-First) Cheat Sheet
-
pydantic.BaseModel— Defines the structure and types for data schemasclass Item(BaseModel): name: str -
@app.get("/path")— Registers a handler for HTTP GET requests@app.get("/health") -
@app.post("/path")— Registers a handler for HTTP POST requests@app.post("/predict") -
response_model=Schema— Enforces the structure and types of the API output@app.post(..., response_model=Result) -
422 Unprocessable Entity— Standard HTTP status for input validation failure -
Field(..., gt=0)— Adds validation constraints (e.g., greater than)value: int = Field(..., gt=0)
Comparison Table
| Endpoint Type | HTTP Method | Primary Use Case |
|---|---|---|
| Health Check | GET | Service status monitoring |
| Prediction | POST | ML inference with payload |
| Configuration Retrieval | GET | Fetching metadata or settings |
Common Pitfalls
- Mistake: Using
dictinstead of Pydantic for request input. Avoid: Always define aBaseModelfor request bodies to get validation. - Mistake: Not setting
response_modelon the route decorator. Avoid: Explicitly define the output schema to guarantee the client contract. - Mistake: Handling validation errors manually inside the route. Avoid: Rely on FastAPI's automatic 422 response handling for consistency.
- Mistake: Putting model loading inside the request handler function. Avoid: Load the model globally or using FastAPI startup events for performance.
FAQs
- Why use Pydantic if I can just parse JSON? Pydantic provides automatic validation, type coercion, documentation, and standardized error handling, saving significant boilerplate code and increasing reliability.
- What is the difference between a 400 and a 422 error? 400 (Bad Request) is generic, often for malformed JSON syntax. 422 (Unprocessable Entity) means the request syntax is correct but the data semantics failed validation against the schema.
- Should I use GET or POST for the /predict endpoint? Use POST. Prediction inputs are typically complex and should be sent in the request body, not exposed or limited by URL query parameters.
- Does FastAPI validate the response body too?
Yes, if you use
response_model, FastAPI validates your function's return value against that schema before serialization, ensuring the client receives the expected structure.