Back to M2 — Experiment Tracking + Registry

Experiment Tracking (MLflow-or-Equivalent)

Outcome: Log params/metrics/artifacts; compare 3 runs Curated video (Ashutosh Tripathi): What is Experiment Tracking in Machine Learning? | MLFlow | — https://www.youtube.com/watch?v=BiUBqv1QNOA (verified live via yt-dlp 2026-09-24). Pointer: tracking harness (starter); shell: courses/video-scripts/mlops-cloud-deploy/03.md.

5 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. Introduction to Tracking — Experiment tracking systems solve the reproducibility crisis by linking code, data, and results.
  2. Logging Parameters and Metrics — Demonstrates initializing an MLflow run and logging scalar inputs and performance outputs.
  3. Handling Artifacts and Seeds — Explains how to store large files like models and plots, and the importance of logging random seeds.
  4. Data Lineage via Checksums — Shows the mechanism for calculating and logging data checksums to verify the training dataset source.
  5. Run Comparison in the UI — Walks through using the MLflow UI to filter, sort, and compare three different training runs side-by-side.
  6. Model Card Generation — Discusses promoting the best run using tags and generating a comprehensive Model Card for registry submission.
PDF notes

Frequently asked questions

What is the difference between a parameter and a metric?

Parameters are inputs that configure the model (e.g., learning rate), while metrics are numerical outputs measuring performance (e.g., accuracy).

Why do we need to log checksums?

Checksums guarantee that the exact training data used for a specific run can be verified later, ensuring data lineage and compliance.

Can I log metrics over time (e.g., epoch loss)?

Yes, `mlflow.log_metric()` supports an optional `step` argument, allowing you to track how a metric evolves during training.

Where does MLflow store the actual model files?

Model files (artifacts) are stored in the configured artifact store, typically an S3 bucket in a cloud deployment, linked by the run ID.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.