Back to From Notebook to Production: The Mental Model

Versioning — Models, Data, Features

When something breaks in production, you need to know exactly which model, which data, which features. Versioning is the audit trail. FIND_VIDEO: search 'model versioning MLflow registry' — recommended channel: MLflow / Made with ML. Aim for 10 min or under.

27 minutesVideo LessonPDF notes
🎯 Free Guest Mode: You are learning for free. Sign in to save your completion progress and quiz answers.

Ready to continue?

Mark this lesson as complete when you're ready to proceed.

Key moments

  1. Tracking Motivation — Unstructured notebook experiments demonstrate why manual printouts fail to maintain provenance across heterogeneous models.
  2. Tracking Server Setup — Installing MLflow and launching the tracking dashboard UI connected to a local SQLite database backend store.
  3. Experiment Configuration — Setting tracking database URIs programmatically in Python and initializing dedicated experiment namespaces.
  4. Manual Run Tracking — Instrumenting a Scikit-Learn run context to log custom tags, hyperparameter dictionaries, evaluation metrics, and model artifacts.
  5. Multi-Model Comparison — Instrumenting CatBoost training and comparing its evaluation metrics against Random Forest within the MLflow UI dashboard.
  6. Deep Learning Workflows — Adapting tracking loops and input structure specifications to log TensorFlow neural network experiments under the same schema.
PDF notes

Frequently asked questions

Why should I use SQLite backend storage over default file storage?

Default file storage writes flat directory files that lack SQL query capabilities. SQLite enables relational indexing, multi-client consistency, and seamless UI filtering.

What happens if the experiment name passed to mlflow.set_experiment does not exist?

MLflow creates the experiment namespace in the database automatically and sets it as the active context for subsequent runs.

Why use framework-specific model loggers instead of standard Python pickle?

Framework loggers package the model alongside conda/pip environment manifests, dependencies, and deserialization flavors required for reproducible downstream serving.

Can I compare Scikit-Learn, CatBoost, and Deep Learning models in one table?

Yes, as long as all runs log to the same experiment namespace using identical metric names like 'rmse'.

How was this lesson?

Your feedback helps us refine explanations and catch bugs.