Training-serving skew is the production ML bug with the worst signal-to-noise ratio. The model works in development, fails in production, and the cause is rarely the model itself.
The mechanic
In training:
def build_features(row):
return [
row['amount'],
row['merchant_id'].lower(), # lowercase
np.log1p(row['user_history_count']), # log transform
row['timestamp'].astype(int), # epoch
]
In serving:
def build_features_v2(req):
return [
req.amount,
req.merchant_id, # not lowercased — bug
req.user_history_count, # not log-transformed — bug
int(datetime.fromisoformat(req.timestamp).timestamp()), # different formula
]
Three subtle differences. The model receives different distributions than it was trained on. Predictions degrade silently.
This happens in real teams constantly. The training pipeline is in a notebook; the serving pipeline is in a microservice. Both written by hand. They drift.
The fix: one source of truth
Option A: Shared library
# Single file imported by both training and serving
# features/feature_engineering.py
def build_features(row):
return {
'amount': row['amount'],
'merchant_id_lower': row['merchant_id'].lower(),
'user_history_log': np.log1p(row['user_history_count']),
'timestamp_epoch': int(row['timestamp'].timestamp()),
}
Both training scripts and serving service import this function. Same logic guaranteed.
Lightweight, no infrastructure. Works for small projects.
Option B: Feature store
A dedicated system that:
- Centralizes feature definitions.
- Computes features in batch (for training).
- Serves features in real-time (for serving).
- Maintains consistency between the two.
Tools: Feast (open-source), Tecton (managed), Hopsworks, AWS SageMaker Feature Store.
What a feature store actually does
For each feature you define:
- Batch processing: compute the feature from historical data for training.
- Online serving: compute (or look up) the feature at prediction time.
- Consistency guarantee: both produce the same value for the same entity at the same point in time.
Example with Feast:
# Define a feature view
from feast import FeatureView, Field
from datetime import timedelta
user_features = FeatureView(
name='user_features',
entities=['user_id'],
ttl=timedelta(days=7),
schema=[
Field('total_orders', Int64),
Field('avg_order_value', Float),
],
source=BigQuerySource(table='user_features_table'),
)
Then:
# Training: get historical features
historical_features = feature_store.get_historical_features(
entity_df=user_entities, # who and when
features=['user_features:total_orders', 'user_features:avg_order_value'],
).to_df()
# Serving: get online features
online_features = feature_store.get_online_features(
entity_rows=[{'user_id': 'cust_123'}],
features=['user_features:total_orders', 'user_features:avg_order_value'],
).to_dict()
Same feature view, same definition, same logic.
Point-in-time correctness
The hardest feature store guarantee: at training time, features must reflect what was known at the time of the prediction event.
Example: predicting churn on 2025-04-15. Features at that point should include "number of logins in last 30 days" computed as of 2025-04-15 — not including events after 2025-04-15.
This is called point-in-time correctness or temporal consistency. Skipping it = subtle training-serving skew.
Feature stores handle this with timestamps:
historical_features = feature_store.get_historical_features(
entity_df=pd.DataFrame({
'user_id': ['cust_123', 'cust_456'],
'event_timestamp': [pd.Timestamp('2025-04-15'), pd.Timestamp('2025-04-15')],
}),
features=[...],
)
Feast computes each feature AS OF the event_timestamp. No future leak.
When you don't need a feature store
For most early-stage ML teams:
- One or two models.
- Features computed from a single source (warehouse).
- Both training and serving query the same warehouse.
A shared library is enough. Pay the feature-store cost (setup, ops) only when you have multiple models sharing features.
When you do need one
- 5+ models sharing features.
- Real-time prediction with sub-100ms latency (warehouse queries too slow).
- Features computed from multiple sources.
- Strict point-in-time requirements.
- Multiple teams collaborating.
At this scale, a feature store amortizes its cost.
The middle ground: SQL-defined features
A pragmatic alternative:
- Features defined in dbt models.
- Training pulls from the dbt model.
- Serving queries the same dbt model (or a materialized table refreshed from it).
Same code, same data, lightweight, no new tool to learn.
Works well for batch-prediction systems and slow-real-time (under 1s latency budget). Not for sub-100ms real-time.
Common training-serving skew bugs
- Different case for strings — 'Email' vs 'email'.
- Different time zones — UTC vs local.
- Different missing-value treatment — NaN vs 0 vs mean imputed.
- Different categorical encoders — train encodes 'unknown' as 0, serving as -1.
- Subtly different aggregations —
last_30_days= 30 days OR 720 hours? Off-by-one mismatches. - Out-of-order features — same names, different column order → wrong predictions.
Audit features at deployment. Compare training-distribution stats to serving-time stats. Discrepancies indicate skew.
Takeaway
Training-serving skew is the #1 production ML bug. Fix it with one source of truth for feature computation — shared library (lightweight) or feature store (at scale). Ensure point-in-time correctness during training. Audit feature distributions at deployment. Without this discipline, your model lies in production.