A real ML service has structure. The seven elements below distinguish a toy from a production system.
Element 1: Request validation
Define the input schema with Pydantic. Reject malformed requests early.
from pydantic import BaseModel, Field
from datetime import datetime
class TransactionRequest(BaseModel):
transaction_id: str = Field(..., min_length=1, max_length=64)
amount: float = Field(..., gt=0, le=1e8)
merchant_id: str = Field(..., min_length=1)
user_id: str = Field(..., min_length=1)
payment_method: str
timestamp: datetime
If a request comes in with amount: -5, FastAPI returns a 422 with a clear error before the model is touched. Saves debugging time.
Element 2: Health and readiness endpoints
@app.get('/health')
def health():
return {'status': 'ok'}
@app.get('/ready')
def ready():
if model is None or feature_store_connection is None:
return {'status': 'not_ready'}, 503
return {'status': 'ready'}
/health for liveness probes (is the service alive). /ready for readiness probes (can it actually serve requests). Kubernetes / Docker orchestrators rely on these to manage your service.
Element 3: Model loading at startup, not per-request
from fastapi import FastAPI
import joblib
app = FastAPI()
model = None
@app.on_event('startup')
def load_model():
global model
model = joblib.load('fraud_model.pkl')
print('Model loaded.')
Loading takes time (potentially seconds). Doing it per-request adds that to every prediction. Load once at startup.
Element 4: Single source of feature computation
Import the same build_features used in training:
from features.feature_engineering import build_features
@app.post('/predict')
def predict(req: TransactionRequest):
features = build_features(req.dict())
prediction = model.predict_proba([features])[0][1]
return {'fraud_probability': float(prediction)}
build_features is the same function used during training-data construction. No drift.
Element 5: Logging every prediction
import logging
logger = logging.getLogger(__name__)
@app.post('/predict')
def predict(req: TransactionRequest):
features = build_features(req.dict())
prediction = model.predict_proba([features])[0][1]
# Log to warehouse for monitoring
log_prediction({
'transaction_id': req.transaction_id,
'features': features,
'prediction': float(prediction),
'model_version': MODEL_VERSION,
'timestamp': datetime.utcnow().isoformat(),
})
return {'fraud_probability': float(prediction)}
Every prediction logged: request, output, model version, timestamp. Required for monitoring and incident debugging.
For latency-sensitive services, log asynchronously (background task) so logging doesn't slow predictions.
Element 6: Graceful error handling
from fastapi import HTTPException
@app.post('/predict')
def predict(req: TransactionRequest):
try:
features = build_features(req.dict())
except FeatureBuildError as e:
logger.error(f'Feature build failed: {e}')
raise HTTPException(status_code=400, detail='Invalid input data')
try:
prediction = model.predict_proba([features])[0][1]
except Exception as e:
logger.error(f'Prediction failed: {e}')
# Fallback: return baseline rate or graceful degradation
return {
'fraud_probability': BASELINE_FRAUD_RATE,
'model_version': 'fallback',
'note': 'Model unavailable, returning baseline'
}
return {'fraud_probability': float(prediction)}
Don't crash the service when one prediction fails. Return a sensible fallback. Log the error for investigation.
Element 7: Versioning in the response
return {
'transaction_id': req.transaction_id,
'fraud_probability': float(prediction),
'model_version': MODEL_VERSION,
'feature_version': FEATURE_VERSION,
'served_at': datetime.utcnow().isoformat(),
}
The consumer (and you, debugging later) knows which model produced the prediction. Critical for A/B testing and audit.
The full minimum service
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
from datetime import datetime
import joblib
import logging
from features.feature_engineering import build_features
from monitoring.logger import log_prediction
logger = logging.getLogger(__name__)
MODEL_VERSION = 'v1.2.0'
FEATURE_VERSION = 'v1.0.0'
BASELINE_FRAUD_RATE = 0.003
app = FastAPI()
model = None
class TransactionRequest(BaseModel):
transaction_id: str
amount: float = Field(..., gt=0)
merchant_id: str
user_id: str
payment_method: str
timestamp: datetime
class PredictionResponse(BaseModel):
transaction_id: str
fraud_probability: float
model_version: str
feature_version: str
served_at: datetime
@app.on_event('startup')
def load_model():
global model
model = joblib.load('fraud_model.pkl')
logger.info(f'Model {MODEL_VERSION} loaded.')
@app.get('/health')
def health():
return {'status': 'ok'}
@app.get('/ready')
def ready():
if model is None:
raise HTTPException(status_code=503, detail='Model not loaded')
return {'status': 'ready'}
@app.post('/predict', response_model=PredictionResponse)
def predict(req: TransactionRequest):
try:
features = build_features(req.dict())
prediction = model.predict_proba([features])[0][1]
except Exception as e:
logger.exception('Prediction failed')
return PredictionResponse(
transaction_id=req.transaction_id,
fraud_probability=BASELINE_FRAUD_RATE,
model_version='fallback',
feature_version=FEATURE_VERSION,
served_at=datetime.utcnow(),
)
log_prediction({
'transaction_id': req.transaction_id,
'prediction': prediction,
'model_version': MODEL_VERSION,
})
return PredictionResponse(
transaction_id=req.transaction_id,
fraud_probability=float(prediction),
model_version=MODEL_VERSION,
feature_version=FEATURE_VERSION,
served_at=datetime.utcnow(),
)
if __name__ == '__main__':
import uvicorn
uvicorn.run(app, host='0.0.0.0', port=8000)
~80 lines. Every element from above. Production-grade.
Testing the service
Before deploying:
# tests/test_predict.py
from fastapi.testclient import TestClient
from app import app
client = TestClient(app)
def test_predict_valid_request():
response = client.post('/predict', json={
'transaction_id': 'test_1',
'amount': 100.0,
'merchant_id': 'merch_a',
'user_id': 'user_1',
'payment_method': 'card',
'timestamp': '2026-05-26T14:00:00Z',
})
assert response.status_code == 200
assert 0 <= response.json()['fraud_probability'] <= 1
def test_predict_invalid_amount():
response = client.post('/predict', json={
'transaction_id': 'test_1',
'amount': -100.0,
# ...
})
assert response.status_code == 422 # validation error
Pytest + TestClient covers most service-level tests. Combine with unit tests on the feature pipeline.
Common service-building mistakes
- Loading the model per request — adds seconds of latency to every prediction.
- No request validation — silent failures on bad input.
- No fallback when prediction fails — service crashes propagate to consumers.
- No model version in response — debugging impossible after the fact.
- Synchronous logging — adds latency to every prediction. Use background tasks.
- No health/ready endpoints — Kubernetes can't manage the service properly.
Takeaway
A production ML service has structure: Pydantic validation, model loaded at startup, single source of feature truth, every prediction logged, graceful error handling, version in response, health endpoints. The 80-line FastAPI template above covers it. Below this bar, you don't have a production service.
Deep Dive: Production Probability Calibration (Platt Scaling & Temperature Scaling)
When serving classification models in production, raw outputs returned by models are often uncalibrated scores rather than true empirical probabilities:
- Tree Ensembles (Random Forest / XGBoost): Output probabilities cluster heavily away from 0 and 1, under-estimating extreme risks.
- Deep Neural Networks: Logits activated via standard Softmax/Sigmoid are frequently over-confident (e.g. predicting 0.98 probability on cases with true empirical accuracy of 80%).
Implementing Platt Scaling with Scikit-Learn
from sklearn.calibration import CalibratedClassifierCV
# Fit a sigmoid calibrator (Platt Scaling) over an independent calibration split
calibrated_clf = CalibratedClassifierCV(base_estimator=clf, method='sigmoid', cv='prefit')
calibrated_clf.fit(X_calib, y_calib)
# Production inference now yields well-calibrated probabilities
calibrated_probabilities = calibrated_clf.predict_proba(X_new)[:, 1]
In high-stakes decisions (e.g. loan approvals or fraud blocks), uncalibrated probabilities distort decision thresholds and lead to severe financial mispricing.