The trained model is a file. To deploy, you need a runnable, environment-locked, scalable unit. That's a Docker container.
Why Docker
Without Docker:
- "Works on my machine, fails on server" (different OS, libraries, Python version).
- Manual setup steps that drift over time.
- "I'll just install one more dependency" → breaks production.
With Docker:
- Build once, run anywhere (same OS, libraries, Python version, file structure).
- Image is immutable: deploy v1 today, deploy v2 tomorrow, roll back to v1 if needed.
- One unit of deployment.
Minimum Dockerfile for an ML service
# Use a specific Python version
FROM python:3.11.5-slim
# Set working directory
WORKDIR /app
# Install dependencies first (better layer caching)
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Copy code
COPY . .
# Expose port
EXPOSE 8000
# Health check (Docker uses this to monitor container health)
HEALTHCHECK --interval=30s --timeout=5s \
CMD curl -f http://localhost:8000/health || exit 1
# Run
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "4"]
Build and test locally:
docker build -t fraud-model:v1.2.0 .
docker run -p 8000:8000 fraud-model:v1.2.0
curl -X POST localhost:8000/predict -H 'Content-Type: application/json' -d '{...}'
If it works locally, it'll work in production (or near enough).
Multi-stage builds for smaller images
The default image is several hundred MB. Optimize:
# Build stage: install dependencies
FROM python:3.11.5-slim AS builder
WORKDIR /app
COPY requirements.txt .
RUN pip install --user --no-cache-dir -r requirements.txt
# Runtime stage: copy only what's needed
FROM python:3.11.5-slim
WORKDIR /app
COPY --from=builder /root/.local /root/.local
COPY . .
ENV PATH=/root/.local/bin:$PATH
EXPOSE 8000
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]
Smaller image = faster deploys, less attack surface, less network transfer.
Image registries
Build the image, push to a registry:
- Docker Hub (public or paid private).
- GitHub Container Registry (free, integrated with GitHub).
- AWS ECR, GCP Artifact Registry, Azure ACR (cloud-native).
docker tag fraud-model:v1.2.0 ghcr.io/yourorg/fraud-model:v1.2.0
docker push ghcr.io/yourorg/fraud-model:v1.2.0
Now any environment can pull that exact image.
Running the container in production
Several options, from simplest to most complex:
1. Single VM with Docker
Just run the container on a server:
docker run -d -p 8000:8000 --restart unless-stopped fraud-model:v1.2.0
- a reverse proxy (Nginx) in front for TLS.
Works for low-traffic services. Cheap. Manual scaling.
2. Managed container services
AWS ECS / Fargate, GCP Cloud Run, Azure Container Apps. You provide a container image; they handle scaling, health checks, deployment.
# Fly.io fly.toml or Cloud Run YAML
service: fraud-model
image: ghcr.io/yourorg/fraud-model:v1.2.0
cpu: 1
memory: 2Gi
min_instances: 1
max_instances: 10
target_cpu: 70%
Best balance for small/medium ML teams. No Kubernetes complexity.
3. Kubernetes
Full container orchestration. Required at large scale. Steep learning curve.
# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: fraud-model
spec:
replicas: 3
selector:
matchLabels:
app: fraud-model
template:
metadata:
labels:
app: fraud-model
spec:
containers:
- name: fraud-model
image: ghcr.io/yourorg/fraud-model:v1.2.0
ports:
- containerPort: 8000
readinessProbe:
httpGet:
path: /ready
port: 8000
livenessProbe:
httpGet:
path: /health
port: 8000
resources:
requests:
memory: "1Gi"
cpu: "500m"
limits:
memory: "2Gi"
cpu: "2000m"
Kubernetes if you have the operational expertise OR are already on a Kubernetes platform. Otherwise managed services are better.
CI/CD for ML services
Use GitHub Actions (or similar) to automate:
# .github/workflows/deploy.yml
on:
push:
tags: ['v*']
jobs:
build-and-deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Login to GHCR
run: echo ${{ secrets.GITHUB_TOKEN }} | docker login ghcr.io -u ${{ github.actor }} --password-stdin
- name: Build image
run: docker build -t ghcr.io/yourorg/fraud-model:${{ github.ref_name }} .
- name: Run tests
run: docker run --rm ghcr.io/yourorg/fraud-model:${{ github.ref_name }} pytest
- name: Push image
run: docker push ghcr.io/yourorg/fraud-model:${{ github.ref_name }}
- name: Deploy (Cloud Run example)
run: gcloud run deploy fraud-model --image=ghcr.io/yourorg/fraud-model:${{ github.ref_name }}
Push a git tag like v1.2.0 → image built → tests run → image pushed → deployed.
Loading the model inside the container
Two approaches:
Bake the model into the image
COPY model.pkl /app/model.pkl
Image includes model. Deploy = update model.
Simple but big images; every retrain = new image.
Mount the model at runtime
Image only has the code. Model is mounted from cloud storage:
@app.on_event('startup')
def load_model():
global model
s3.download_file('my-bucket', f'models/v{MODEL_VERSION}/model.pkl', '/tmp/model.pkl')
model = joblib.load('/tmp/model.pkl')
Image is small. Multiple models can share image. New model = update config, restart pod.
Best for production: mount at runtime, version via config.
Common containerization mistakes
pip installwithout version pins — image diverges over rebuilds.- Using
latesttag everywhere — non-reproducible. Pin to specific versions. - Single-stage build with dev deps — huge images.
- No HEALTHCHECK — orchestrator can't manage your container's lifecycle.
- Running as root — security risk. Add a non-root user.
- No resource limits — container can hog CPU/memory.
Takeaway
Docker locks the environment. Multi-stage builds keep images small. Push to a registry. Run on managed container services (Cloud Run, Fargate) for most teams; Kubernetes for scale. Automate via CI/CD. Mount the model at runtime, version via config. This pattern is the lightweight cloud-native default.