Reproducibility = given the same starting point, get the same result. Sounds trivial; in ML, it's surprisingly hard. Four things must be locked.
Lock 1: Code
Obvious: code in version control (git). Less obvious: every script that touched the training pipeline.
- Feature engineering scripts.
- Training code.
- Hyperparameter tuning scripts.
- Validation logic.
If any of these live in a notebook that gets overwritten, or in a "private" branch, you've already lost reproducibility.
Pattern: training is a runnable script with explicit inputs and outputs. Not a notebook for production. Notebooks are for exploration.
python train.py --data-path data/v3 --output-dir models/v1.0.0 --config configs/xgb_v1.yaml
Lock 2: Data
Code + same data = same model only if the data is the same.
For real datasets that change over time, you need data versioning. Three options:
DVC (Data Version Control)
Tracks data files in S3/GCS with git-friendly hashes.
dvc add data/training_data.parquet
git add data/training_data.parquet.dvc data/.gitignore
git commit -m "Training data v3"
dvc push # uploads to remote storage
Now git checkout <commit> + dvc pull reproduces the exact data state.
Date-stamped snapshots
Simpler:
# At training time
training_data = load_data(snapshot_date='2026-05-26')
Snapshot path: s3://data/snapshots/2026-05-26/. Date in code; data immutable in storage.
Commit the data
Only if it's small (<100MB). Most ML datasets are too big for this.
DVC is best practice; date-stamped snapshots are good-enough lightweight.
Lock 3: Environment
Same code + same data + same library versions = same model.
# Bad
pip install scikit-learn xgboost pandas
# Good
pip install scikit-learn==1.4.0 xgboost==2.0.3 pandas==2.1.4
Better: lock the whole environment.
requirements.txt with pinned versions
scikit-learn==1.4.0
xgboost==2.0.3
pandas==2.1.4
numpy==1.26.0
Poetry / pip-tools (lockfile)
Captures transitive dependencies too. More robust.
Docker (best)
Lock the OS, Python version, library versions, all of it.
FROM python:3.11.5-slim
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . /app
WORKDIR /app
CMD ["python", "train.py"]
docker build -t my-ml-trainer:v1 . creates a reproducible environment that runs the same on anyone's machine.
For training: Docker is overkill for solo iteration but worth it for shared/CI training.
For serving: Docker is mandatory. Production environment must match training-time guarantees.
Lock 4: Randomness
ML algorithms often use randomness:
- Train/test splits.
- Bootstrap samples (random forest).
- Initial weights (neural networks).
- Mini-batch sampling.
- Cross-validation fold assignment.
Without seeding, two runs give slightly different models. Seemingly small differences cascade into different hyperparameter selections, different test scores, irreproducible results.
import random
import numpy as np
import torch
random.seed(42)
np.random.seed(42)
torch.manual_seed(42)
torch.cuda.manual_seed_all(42)
torch.backends.cudnn.deterministic = True # for GPU determinism
For sklearn / XGBoost: pass random_state=42 to every estimator and splitter.
X_train, X_test = train_test_split(X, y, random_state=42)
model = XGBClassifier(random_state=42)
Tracking experiments
Even with all four locks, you'll run many experiments. Track which model came from which:
- Data version.
- Code commit.
- Config used.
- Hyperparameters.
- Resulting metrics.
MLflow:
import mlflow
with mlflow.start_run():
mlflow.log_param('learning_rate', 0.05)
mlflow.log_param('max_depth', 6)
mlflow.log_metric('val_auc', 0.85)
mlflow.log_metric('test_auc', 0.83)
mlflow.sklearn.log_model(model, 'model')
mlflow.log_artifact('config.yaml')
mlflow.set_tag('data_version', 'v3')
mlflow.set_tag('git_commit', subprocess.check_output(['git', 'rev-parse', 'HEAD']).strip())
Now mlflow ui shows every run. Re-find the best one. Compare params. Reproduce by checking out the git commit, restoring the data version, restoring the environment.
What "reproducible" actually means
Two flavors:
Bit-for-bit reproducible
Same code + same data + same env + same seed = exactly the same model.
Possible for CPU-only training. Hard with GPU (non-deterministic kernels, atomic ops).
Statistically reproducible
Same general setup → models with very similar performance, even if exact weights differ.
This is what you usually need. Match validation metrics within 0.5%; cluster centroids within 1%.
For most ML systems, statistical reproducibility is the goal. Don't over-engineer for bit-for-bit unless you have a regulatory reason.
Common reproducibility mistakes
- Notebooks as training code — cell execution order matters; state leaks across cells.
pip installwithout versions — broken in 6 months when libraries update.- Not seeding — train/test splits, model init randomness silently shifts results.
- Mutable data sources — training on "yesterday's snapshot" without snapshotting.
- Local environment differs from CI — works on laptop, fails on the cluster.
Takeaway
Reproducibility requires locking four things: code (git), data (DVC or snapshots), environment (pinned versions + Docker), and seed (everywhere). Track experiments with MLflow. Aim for statistical reproducibility; bit-for-bit only when regulated. Without these locks, your "best model" is unrecoverable in 6 months.