This lesson on Experiment Tracking (MLflow-or-Equivalent) is hands-on and example-driven. You will be able to systematically record all inputs, outputs, and metadata for machine learning training runs using an experiment tracking system like MLflow. You will learn how to log parameters, metrics, and large artifacts, enabling you to compare multiple runs side-by-side to determine the optimal model. Finally, you will generate a comprehensive model card for the selected run, preparing it for deployment via the Model Registry.
What You'll Be Able To Do
- Initialize and configure an MLflow tracking session for a training script.
- Log scalar parameters, performance metrics, and random seeds associated with a training run.
- Store large model files and data checksums as trackable run artifacts.
- Utilize the tracking UI to compare the metrics and parameters of three distinct experiment runs.
- Generate a structured Model Card summarizing the performance and lineage of a promoted run.
- Integrate experiment tracking into an existing M1 serving contract workflow.
Detailed Concept Walkthrough
1. Core Tracking Components
Experiment tracking links code, data, and results, providing reproducibility. MLflow organizes runs within named experiments, ensuring every training iteration is logged and traceable.
- Mechanism: MLflow uses a backend store (file system, database) to record metadata (params, metrics) and an artifact store (S3, local path) for large files (models, plots). This separation optimizes storage and retrieval by keeping small, searchable data separate from large binary objects.
- Execution Flow: A run starts with
mlflow.start_run(), which generates a unique Run ID. All subsequent logging calls within the context manager are automatically associated with this ID, ensuring atomic logging and preventing data loss if the script crashes. - Best Practice: Always define an experiment name using
mlflow.set_experiment()before starting runs; this groups related iterations logically and prevents logging to the default experiment, which quickly becomes cluttered and unmanageable.
import mlflow
import numpy as np
# Set the experiment name for grouping related runs
mlflow.set_experiment("Churn_Prediction_V2")
with mlflow.start_run(run_name="Linear_Model_Trial"):
# 1. Log parameters (inputs)
learning_rate = 0.01
epochs = 10
mlflow.log_param("learning_rate", learning_rate)
mlflow.log_params({"epochs": epochs, "seed": 42})
# 2. Simulate training and log metrics (outputs)
loss = np.random.rand() * 0.5
accuracy = 0.85 + np.random.rand() * 0.1
mlflow.log_metric("validation_loss", loss)
mlflow.log_metric("accuracy", accuracy, step=epochs)
print(f"MLflow Run ID: {mlflow.active_run().info.run_id}")
Key Takeaway: Tracking ensures every run is uniquely identified and its configuration (parameters) and results (metrics) are permanently recorded.
2. Managing Large Artifacts
Artifacts are large output files—like trained models, plots, or data samples—that cannot be stored efficiently in the metadata database. They are stored separately in the artifact store and linked via the run ID.
- Mechanism:
mlflow.log_artifact()copies a local file into the designated artifact storage location (e.g., S3 bucket or local directory) under the run's path. This process ensures the artifact is immutable and tied specifically to the run that produced it, guaranteeing version control. - Under the Hood: When logging a model, MLflow often uses flavor functions like
mlflow.sklearn.log_model(), which saves the model object and automatically includes environment dependencies (Conda/pip) required for later loading. This simplifies the deployment step by bundling necessary environment context. - Best Practice: Log data checksums (e.g., MD5 or SHA256) as parameters or artifacts alongside the model. This guarantees that the exact dataset used for training can be verified later, fulfilling the critical lineage requirement for regulatory compliance and debugging.
import mlflow
import hashlib
# Assume 'model.pkl' and 'train_data.csv' exist locally
def calculate_checksum(filepath):
"""Calculates SHA256 checksum for a file."""
hash_sha256 = hashlib.sha256()
with open(filepath, "rb") as f:
for chunk in iter(lambda: f.read(4096), b""):
hash_sha256.update(chunk)
return hash_sha256.hexdigest()
with mlflow.start_run():
# Log the trained model file
mlflow.log_artifact("model.pkl", artifact_path="models")
# Calculate and log the checksum of the training data file
data_path = "train_data.csv"
data_hash = calculate_checksum(data_path)
mlflow.log_param("training_data_checksum_sha256", data_hash)
Key Takeaway: Artifacts are immutable files stored externally, and logging checksums ensures data lineage and reproducibility.
3. Comparison and Model Cards
The tracking UI allows side-by-side comparison of runs based on logged parameters and metrics, facilitating the selection of the best model. The Model Card formalizes this selection for promotion.
- Mechanism: The MLflow UI queries the backend store to retrieve all runs within an experiment. It then displays metrics (like AUC or F1) and parameters (like learning rate) in sortable columns, allowing rapid identification of high-performing configurations based on business criteria.
- Best Practice: When comparing runs, focus on maximizing the primary business metric (e.g., minimizing false negatives) rather than just maximizing accuracy. Use tags (
mlflow.set_tag()) to categorize runs (e.g., 'production_candidate', 'hyperparameter_search') for easier filtering and promotion. - Execution Flow: Once a run is selected, a Model Card is generated. This document aggregates all relevant metadata (code version, training data checksum, performance metrics, intended use) into a single, human-readable report that serves as the input for the Model Registry (M2-L8).
import mlflow
# After identifying the best run ID (e.g., via UI comparison)
BEST_RUN_ID = "a1b2c3d4e5f6"
# 1. Set a tag to mark the run for promotion
with mlflow.start_run(run_id=BEST_RUN_ID):
mlflow.set_tag("promotion_status", "Candidate_M2L8")
mlflow.set_tag("reviewer", "Data_Scientist_A")
# 2. Retrieve data and generate Model Card content
run_data = mlflow.get_run(BEST_RUN_ID).data
model_card_content = f"""
# Model Card for Run {BEST_RUN_ID}
## Performance Summary
- F1 Score: {run_data.metrics.get('f1_score', 'N/A')}
- Training Data Checksum: {run_data.params.get('training_data_checksum_sha256', 'N/A')}
"""
with open("model_card.md", "w") as f:
f.write(model_card_content)
# 3. Log the final Model Card as an artifact
with mlflow.start_run(run_id=BEST_RUN_ID):
mlflow.log_artifact("model_card.md")
Key Takeaway: Use the tracking UI to compare metrics across runs, and formalize the selection by tagging the winner and generating a comprehensive Model Card artifact.
Topics Covered in Experiment Tracking (MLflow-or-Equivalent)
- Introduction to Tracking (0:00 - 1:15) — Experiment tracking systems solve the reproducibility crisis by linking code, data, and results.
- Logging Parameters and Metrics (1:15 - 3:00) — Demonstrates initializing an MLflow run and logging scalar inputs and performance outputs.
- Handling Artifacts and Seeds (3:00 - 4:45) — Explains how to store large files like models and plots, and the importance of logging random seeds.
- Data Lineage via Checksums (4:45 - 6:30) — Shows the mechanism for calculating and logging data checksums to verify the training dataset source.
- Run Comparison in the UI (6:30 - 8:00) — Walks through using the MLflow UI to filter, sort, and compare three different training runs side-by-side.
- Model Card Generation (8:00 - 9:00) — Discusses promoting the best run using tags and generating a comprehensive Model Card for registry submission.
MLOps + Cloud Deploy (AWS-First) Cheat Sheet
-
mlflow.set_experiment("name")— Groups related runs under a common namemlflow.set_experiment("Hyperparam_Search_V3") -
mlflow.log_param("key", value)— Records a single input setting or configurationmlflow.log_param("max_depth", 10) -
mlflow.log_metric("key", value, step=N)— Records a single numerical performance resultmlflow.log_metric("f1_score", 0.915, step=5) -
mlflow.log_artifact("path", "folder")— Uploads a local file to the run's artifact storemlflow.log_artifact("model.pkl", "models/") -
mlflow.set_tag("key", value)— Adds descriptive, searchable metadata to a runmlflow.set_tag("data_version", "2023-Q4") -
Checksum Logging— Verifies data integrity and lineagemlflow.log_param("data_hash", data_md5)
Comparison Table
| Feature | Purpose | Data Type | Storage Location |
|---|---|---|---|
| Parameters | Input configuration (static). | String/Number | Backend Store (DB/File) |
| Metrics | Output performance (numerical). | Number (float) | Backend Store (DB/File) |
| Artifacts | Large files (models, plots). | Binary/File | Artifact Store (S3/Local) |
| Tags | Descriptive run metadata (searchable). | String | Backend Store (DB/File) |
Common Pitfalls
- Mistake: Logging metrics outside the active run context.
Avoid: Always wrap logging calls within
with mlflow.start_run():. - Mistake: Storing large model files directly as parameters.
Avoid: Use
mlflow.log_artifact()or flavor-specificlog_model()functions. - Mistake: Forgetting to log the random seed used for training. Avoid: Log the seed as a parameter to ensure run reproducibility.
- Mistake: Using the default experiment name for production runs.
Avoid: Explicitly call
mlflow.set_experiment()before starting any run.
FAQs
- What is the difference between a parameter and a metric? Parameters are inputs that configure the model (e.g., learning rate), while metrics are numerical outputs measuring performance (e.g., accuracy).
- Why do we need to log checksums? Checksums guarantee that the exact training data used for a specific run can be verified later, ensuring data lineage and compliance.
- Can I log metrics over time (e.g., epoch loss)?
Yes,
mlflow.log_metric()supports an optionalstepargument, allowing you to track how a metric evolves during training. - Where does MLflow store the actual model files? Model files (artifacts) are stored in the configured artifact store, typically an S3 bucket in a cloud deployment, linked by the run ID.