This lesson on Versioning — Models, Data, Features is hands-on and example-driven. You will set up a persistent MLflow tracking server backed by SQLite, instrument machine learning code to log hyperparameters, metrics, and serialized artifacts, and compare heterogeneous models in the MLflow UI. This workflow establishes reproducible experiment tracking across Scikit-Learn, CatBoost, and deep learning architectures.
What You'll Be Able To Do
- Launch an MLflow tracking server using an SQLite metadata store via the CLI.
- Configure tracking URIs and experiment namespaces programmatically in Python.
- Log hyperparameters, custom metadata tags, evaluation metrics, and serialized artifacts inside run contexts.
- Compare cross-architecture model runs and performance metrics within the MLflow UI dashboard.
- Instrument deep learning pipelines with structured input specifications and framework-specific model loggers.
Detailed Concept Walkthrough
1. Experiment Tracking Architecture and Server Setup
MLflow decouples the web user interface from the underlying metadata database and artifact repository. Running a dedicated tracking backend enables structured relational querying rather than fragile file-system dumps.
- Mechanism: Launching the MLflow UI with an explicit backend store URI directs all run metadata, metrics, and parameters into a structured SQLite database. This ensures complete auditability and multi-run persistence across notebook sessions.
- Under the Hood: MLflow writes scalar values, runtime tags, and parameter dictionaries directly to relational tables in the SQLite database file (
mlflow.db), while heavy binary artifacts like serialized models are persisted to an artifact directory. - Best Practice: Omitting
--backend-store-uriforces MLflow to fall back to raw file-based directory storage (./mlruns). This fallback mode lacks robust relational querying and frequently causes synchronization mismatches between code runs and UI views.
# Shell: Start the MLflow tracking server with SQLite backend store
mlflow ui --backend-store-uri sqlite:///mlflow.db --port 5000
Key Takeaway: Always specify an explicit database backend URI when launching MLflow to ensure relational querying and avoid file synchronization issues.
2. Programmatic Experiment Binding and Namespaces
Experiments group related model iterations under a single conceptual namespace to isolate distinct business problems or dataset targets. Binding code to the correct tracking URI ensures all subsequent logs land in the designated backend.
- Mechanism: Setting the tracking URI in your Python environment establishes the connection to the SQLite database before any training begins. Initializing an experiment namespace sets the active target for all subsequent run executions.
- Execution Flow:
mlflow.set_experiment()checks if the specified experiment namespace already exists in the backend database. If the experiment is absent, MLflow automatically creates it and sets it as the active context for child runs. - Best Practice: Separate exploratory modeling efforts into clearly named experiments rather than logging everything to the default 'Default' experiment (ID
0). This practice keeps project metrics clean and prevents metric cross-contamination.
import mlflow
# Configure connection to backend SQLite store
mlflow.set_tracking_uri("sqlite:///mlflow.db")
# Set or auto-create project-specific experiment namespace
mlflow.set_experiment("income_prediction")
Key Takeaway: Define explicit tracking URIs and dedicated experiment namespaces at the entry point of every training pipeline.
3. Context-Managed Run Logging and Model Serialization
The mlflow.start_run() context manager encapsulates an isolated training iteration, ensuring all parameters, tags, metrics, and model binaries are atomically linked to a unique run ID.
- Mechanism: Within the run context block,
log_params()captures configuration dictionaries,set_tag()attaches operational metadata like model family, andlog_metric()logs final evaluation scores. - Under the Hood: Framework-specific model loggers (such as
mlflow.sklearn.log_model) serialize the model object alongside an MLmodel metadata file, Python environment specifications (conda.yaml,requirements.txt), and deserialization flavors. - Syntax Rule: Always use framework-specific loggers over generic serializers like raw
pickle. Framework loggers bundle required dependency manifests, allowing downstream serving tools to accurately reproduce the execution environment.
import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import root_mean_squared_error
params = {"n_estimators": 100, "max_depth": 6, "random_state": 42}
with mlflow.start_run(run_name="rf_baseline"):
# Log operational metadata and hyperparameters
mlflow.set_tag("model_family", "random_forest")
mlflow.log_params(params)
# Train and evaluate
rf = RandomForestRegressor(**params)
rf.fit(X_train, y_train)
rmse = root_mean_squared_error(y_val, rf.predict(X_val))
# Log evaluation metrics and framework-specific model artifact
mlflow.log_metric("rmse", rmse)
mlflow.sklearn.log_model(rf, artifact_path="model")
Key Takeaway: Wrap model training in run contexts and log artifacts using framework-specific loggers to package dependencies alongside weights.
4. Multi-Model Benchmarking Across Heterogeneous Frameworks
Standardizing metric keys and metadata tags across disparate algorithms allows direct, unbiased performance comparisons inside a centralized experiment registry.
- Mechanism: Different algorithms—such as CatBoost, Scikit-Learn ensembles, and deep neural networks—are trained within the same experiment namespace using identical metric names (e.g.,
rmse). - Execution Flow: For deep learning libraries like TensorFlow, models require structured input layer definitions before logging. Once trained, framework loggers (
mlflow.tensorflow.log_modelormlflow.catboost.log_model) package the model under identical experiment schemas. - Best Practice: Maintain strict consistency in metric naming conventions across all framework scripts. Mismatched metric keys (e.g.,
RMSEvsrmsevsval_loss) prevent the UI from generating unified tabular comparisons and scatter plots.
from catboost import CatBoostRegressor
import mlflow
cb_params = {"iterations": 500, "learning_rate": 0.05, "depth": 6}
with mlflow.start_run(run_name="catboost_tuned"):
mlflow.set_tag("model_family", "catboost")
mlflow.log_params(cb_params)
cb = CatBoostRegressor(**cb_params, verbose=0)
cb.fit(X_train, y_train)
val_rmse = root_mean_squared_error(y_val, cb.predict(X_val))
# Consistent metric name enables multi-model UI comparison
mlflow.log_metric("rmse", val_rmse)
# Persist via CatBoost or generic artifact logger
Key Takeaway: Enforce unified metric names across heterogeneous algorithms to enable instant cross-model benchmarking in the tracking UI.
Topics Covered in Versioning — Models, Data, Features
- Tracking Motivation (0:00 - 2:49) — Unstructured notebook experiments demonstrate why manual printouts fail to maintain provenance across heterogeneous models.
- Tracking Server Setup (2:49 - 5:05) — Installing MLflow and launching the tracking dashboard UI connected to a local SQLite database backend store.
- Experiment Configuration (5:05 - 6:51) — Setting tracking database URIs programmatically in Python and initializing dedicated experiment namespaces.
- Manual Run Tracking (6:51 - 10:11) — Instrumenting a Scikit-Learn run context to log custom tags, hyperparameter dictionaries, evaluation metrics, and model artifacts.
- Multi-Model Comparison (10:11 - 12:00) — Instrumenting CatBoost training and comparing its evaluation metrics against Random Forest within the MLflow UI dashboard.
- Deep Learning Workflows (12:00 - 13:00) — Adapting tracking loops and input structure specifications to log TensorFlow neural network experiments under the same schema.
ML in Practice Cheat Sheet
-
mlflow ui --backend-store-uri <uri>— Launch UI dashboard connected to specified metadata storemlflow ui --backend-store-uri sqlite:///mlflow.db -
mlflow.set_tracking_uri(uri)— Configure backend tracking database URI in Pythonmlflow.set_tracking_uri("sqlite:///mlflow.db") -
mlflow.set_experiment(experiment_name)— Set active experiment namespace or create if missingmlflow.set_experiment("income") -
with mlflow.start_run(run_name=...):— Initialize context manager for isolated experiment runwith mlflow.start_run(run_name="rf_depth_10"): -
mlflow.set_tag(key, value)— Attach custom string metadata or identifiers to runmlflow.set_tag("model_type", "catboost") -
mlflow.log_params(params_dict)— Log dictionary of hyperparameter key-value pairsmlflow.log_params({"lr": 0.01, "epochs": 50}) -
mlflow.log_metric(key, value)— Record single scalar evaluation metric valuemlflow.log_metric("rmse", 0.2451) -
mlflow.sklearn.log_model(sk_model, artifact_path)— Serialize Scikit-Learn model with dependency flavorsmlflow.sklearn.log_model(rf, "model")
Comparison Table
| Dimension | Manual Logging (log_*) | Automatic Logging (autolog) |
|---|---|---|
| Granularity | Explicit control over exact logged parameters | Automatically captures all library parameter defaults |
| Artifact Customization | Requires explicit model and artifact calls | Auto-saves standard model binaries and summaries |
| Metadata Tagging | Enables custom identifiers and operational tags | Logs only standard framework-level metadata tags |
| Code Footprint | Requires several lines of instrumentation code | Requires single mlflow.autolog() invocation call |
Common Pitfalls
- Mistake: Launching MLflow UI without specifying the SQLite backend store URI. Avoid: Always pass the backend store URI flag pointing directly to your SQLite database file.
- Mistake: Persisting models using raw pickle dumps instead of framework loggers. Avoid: Use framework-specific methods like mlflow.sklearn.log_model to package required dependency specifications.
- Mistake: Running training iterations without declaring an experiment namespace. Avoid: Call mlflow.set_experiment before executing runs to group iterations cleanly away from the default namespace.
- Mistake: Using inconsistent metric key names across different model architectures. Avoid: Standardize metric naming across all model scripts to enable direct comparison inside the UI dashboard.
FAQs
- Why should I use SQLite backend storage over default file storage? Default file storage writes flat directory files that lack SQL query capabilities. SQLite enables relational indexing, multi-client consistency, and seamless UI filtering.
- What happens if the experiment name passed to mlflow.set_experiment does not exist? MLflow creates the experiment namespace in the database automatically and sets it as the active context for subsequent runs.
- Why use framework-specific model loggers instead of standard Python pickle? Framework loggers package the model alongside conda/pip environment manifests, dependencies, and deserialization flavors required for reproducible downstream serving.
- Can I compare Scikit-Learn, CatBoost, and Deep Learning models in one table? Yes, as long as all runs log to the same experiment namespace using identical metric names like 'rmse'.