Micro Free Course — 100% Free Learning

All lessons in this module are free to learn. Sign in with Google to save your progress.

1

Serving and Deployment

Batch vs online inference, building a FastAPI prediction service, Docker / Kubernetes packaging, scaling and latency budgets, A/B testing and champion-challenger model rollouts.

Module Progress0% Complete
294 min total
12 Lessons
0 Completed

Module Content

Batch vs Online Inference

Two serving patterns. The wrong choice causes 90% of avoidable infrastructure pain. FIND_VIDEO: search 'batch vs online inference machine learning' — recommended channel: Made with ML / Chip Huyen. Aim for 10 min or under.

4 minVideo
Start

Quiz: Choosing Your Serving Pattern

Batch is cheap and easy. Online is hard and necessary only sometimes. Pick deliberately.

7 minTutorial
Start

Building a Prediction Service — The FastAPI Pattern

FastAPI + a Pydantic schema + a loaded model = production-grade prediction service. The standard pattern. FIND_VIDEO: search 'fastapi machine learning model serving' — recommended channel: FastAPI / Made with ML. Aim for 11 min or under.

18 minVideo
Start

Quiz: Anatomy of a Production ML Service

The seven elements of a real ML service. Without each, you have a toy.

7 minTutorial
Start

Containerization and Deployment (Docker, Kubernetes)

Docker locks the environment; Kubernetes (or simpler alternatives) runs the container. The lightweight version of cloud-native ML. FIND_VIDEO: search 'docker machine learning deployment' — recommended channel: Made with ML / Docker docs. Aim for 11 min or under.

21 minVideo
Start

Quiz: Packaging Models for Deployment

A model in a pickle file is unshippable. A Docker container is shippable. The packaging layer that connects training to serving.

7 minTutorial
Start

Case 2 — Deploy a Model as a Containerized FastAPI Service

Package the registered churn model into a production-ready, low-latency FastAPI inference service and optimize it within a multi-stage Docker container.

103 minSubmission
Start

Scaling — Autoscaling, Latency Budgets, Caching

How to keep predictions fast and the bill manageable as traffic grows. FIND_VIDEO: search 'ML service latency caching autoscaling' — recommended channel: Made with ML / Chip Huyen. Aim for 10 min or under.

13 minVideo
Start

Quiz: Latency Budgets and Optimization

Latency is a contract. Setting one explicitly and optimizing for it produces a system you can actually run.

7 minTutorial
Start

A/B Testing and Champion-Challenger

The discipline of deploying new models. Never swap; always test. FIND_VIDEO: search 'ml model ab test champion challenger' — recommended channel: Statsig / Made with ML. Aim for 10 min or under.

7 minVideo
Start

Quiz: Safely Rolling Out New Models

The three-step rollout: shadow → canary → full. Each step catches a different class of bug before it hits all users.

7 minTutorial
Start

Case 3 — Run an A/B Test Between Two Model Versions

Design and simulate an enterprise A/B champion-challenger canary deployment with hash-based routing, SRM checks, and statistical decision rules.

93 minSubmission
Start