Join the AI + Data Tour for hands-on training, real customer stories, and time with Domo product experts near you.
Machine Learning Model Management: How to Build, Train, and Monitor ML Models

Getting a machine learning model from a Jupyter notebook to production is one challenge. Keeping it accurate, governed, and maintainable over time? That's where most teams stumble. This guide walks through the complete ML model lifecycle, explains the difference between model management and machine learning operations (MLOps), and provides practical frameworks for versioning, monitoring, and governance that enterprise teams need to operationalize machine learning at scale.
Key takeaways
Here are the main points to keep in mind:
- ML model management encompasses the full lifecycle of building, deploying, monitoring, and governing machine learning models in production environments
- The seven stages of the ML model lifecycle span from data collection through ongoing monitoring and maintenance
- Five core components form the foundation of effective model management: versioning, code versioning, experiment tracking, model registry, and model monitoring
- Governance and compliance are essential for enterprise ML, ensuring auditability, reproducibility, and responsible AI practices
- Successful model management requires automation, clear metrics, and continuous monitoring to prevent model degradation
What are machine learning models?
Machine learning models are algorithms trained on data to recognize patterns and make predictions or decisions. The goal of machine learning is to train systems to recognize patterns and make predictions or decisions from data. Programmers want the program to be able to make its own decisions.
ML is built on foundations of data, but what makes ML different from other parts of computer science is the way it processes that data. Much like the way you learn a foreign language, the ML model needs practice to "learn" how to interpret the data it intakes and give accurate outputs.
To make sure the patterns are accurate, programmers feed data into ML programs many times. The part of machine learning that actually processes the data is called the ML model. To increase model accuracy, programmers may change the way the program processes data (the model architecture). Programmers train the algorithm on a set of data, and if the output is correct, they know the model works. If the result is off, though, they know they need to continue the process of adjusting model architecture, evaluating model inputs, and retraining the model on new data.
At this point, ML is an important capability for competitive organizations. Organizations that implement ML can gain value from task automation, predictive maintenance, use case personalization, and predictive modeling.
4 types of machine learning models
Four kinds of machine learning models exist, each suited to different problems and data scenarios:
- Supervised learning models are models that are given labeled data. The model is fed a set of data; if the data has been cleaned and labeled correctly, programmers can teach the model to evaluate and recognize patterns in that data. The model learns from the script it is given. From a management perspective, supervised models require careful governance of labeled training data and clear documentation of labeling criteria.
- Semi-supervised learning models use a small amount of labeled data and a larger amount of unlabeled data. These models learn basic frameworks of pattern recognition from the information about the data they're given and then infer the rest about the unlabeled data. Managing these models means tracking both the labeled subset and the inference logic applied to unlabeled portions.
- Unsupervised learning models need no human interaction. They are unsupervised and free to process unlabeled data on their own. As models process the data, they may notice patterns and create "clusters" of commonalities between data points. Drift detection becomes especially important here since cluster stability can shift as new data arrives.
- Reinforcement learning for models is simply a trial-and-error process. Programmers allow the model to process its data with a certain goal in mind, rewarding the program when it succeeds and disciplining it when it doesn't. The program attempts to process the data in order to meet the goal. The model learns to create algorithms that create outputs that achieve the goal in the most effective way. These models require monitoring of reward signals and careful versioning of reward functions.
What is machine learning model management?
ML model management is the discipline of creating, evaluating, versioning, deploying, and monitoring machine learning models across their entire lifecycle. It sits within the broader MLOps practice but focuses specifically on the model itself rather than the surrounding infrastructure, data pipelines, or organizational processes.
The boundaries matter because adjacent terms often get conflated. Here's how model management relates to related concepts:
A practical way to think about it: if your focus is getting a single model from notebook to production, you're doing model management. Building the platform that enables many teams to do that repeatedly? That's MLOps. Ensuring those models meet regulatory requirements falls under model governance.
With stronger ML models, companies can improve customer experience, communicate more effectively with customers through chatbots, and automate rote tasks to save time and money. According to McKinsey's 2026 State of AI survey, 88 percent of organizations now regularly use AI in at least one business function, with respondents most frequently reporting cost reductions from AI use in supply chain management, service operations, and manufacturing.
Depending on the project's or company's goal, programmers may use different types of models, varying model architectures, a range of inputs, different data sets, and modified parameters. Programmers keep track of the results of each model to help them evaluate which performs the best. When looking at model results, programmers must keep in mind the accuracy, reliability, and efficiency of the model.
Models vs experiments in model management
AI model management needs ongoing configuration experimentation and tracking of the models. Within ML model management, there are two important parts: models and experiments. These two complementary yet distinct parts comprise model management.
The model-focused aspect of managing models encompasses model packaging, lineage, deployment, tracking performance, and then (if necessary) retraining the model based on its performance.
The other piece of model management is experimentation. Machine learning is built on trial-and-error processes, and experimentation is no exception. In this realm of ML model management, programmers can experiment with the model, pushing its parameters, tracking metrics about model performance, gathering data, and pipeline versioning. This data helps programmers and data scientists collaborate on models and look at background info on each model.
The 7 stages of the machine learning model lifecycle
Understanding the full model lifecycle helps teams manage each phase effectively. While the specifics vary by organization, most ML projects move through seven distinct stages:
- Data collection: Gathering raw data from various sources including databases, APIs, sensors, and third-party providers
- Data preparation: Cleaning, transforming, labeling, and splitting data into training, validation, and test sets
- Model engineering: Selecting algorithms, designing architectures, and building the initial model structure
- Model selection: Comparing multiple candidate models to identify the best performer for the task
- Training and validation: Teaching the model on training data and validating performance on held-out data
- Deployment: Moving the trained model into a production environment where it can serve predictions
- Monitoring and maintenance: Tracking model performance over time and retraining when accuracy degrades
These stages map to common ML frameworks, though terminology varies:
Each stage presents its own challenges and requires specific tools and practices. The transitions between stages are where many ML projects struggle, which is why model management practices focus heavily on handoffs, documentation, and automation.
{{custom-cta-1}}
Data collection and preparation
The first two stages establish the foundation for everything that follows. Data collection involves identifying relevant data sources, establishing data pipelines, and ensuring you have sufficient volume and variety for your use case. Strong data collection helps models generalize across a wider range of scenarios.
Data preparation is often the most time-consuming stage. Handling missing values. Removing duplicates. Normalizing formats. Creating meaningful features. For supervised learning, this stage also involves labeling data accurately, which can require significant human effort or specialized annotation tools. This is the stage many teams rush through to get to the "interesting" model work, and the problems that creates will haunt you throughout the lifecycle. Inconsistent labeling or poorly handled missing values will surface later, guaranteed.
Model development and training
Model engineering, selection, and training form the core development phase. During model engineering, data scientists choose appropriate algorithms based on the problem type, whether that's classification, regression, clustering, or something else. They design the model architecture, including decisions about layers, parameters, and optimization approaches.
Model selection involves running experiments with multiple candidate models and comparing their performance. Experiment tracking becomes critical here. Teams need to record which configurations they tried, what hyperparameters they used, and how each variant performed.
Training and validation then refine the selected model. The model learns patterns from training data while validation data helps prevent overfitting. Multiple iterations are typical as teams tune hyperparameters and adjust architectures based on validation results.
Evaluation and deployment
Before deployment, models undergo rigorous evaluation against test data the model has never seen. This provides an unbiased estimate of how the model will perform in production. Evaluation metrics vary by task but might include accuracy, precision, recall, F1 score, or mean squared error.
Deployment moves the model from a development environment into production. Several forms exist depending on the use case. Batch deployment processes large volumes of data on a schedule, while real-time deployment serves predictions on demand with low latency. Some organizations use shadow deployment to run new models alongside existing ones before fully switching over.
Monitoring and maintenance
Once deployed, models require ongoing attention. Model monitoring tracks performance metrics, watches for data drift, and alerts teams when something goes wrong. With monitoring, teams can catch performance drift as production data diverges from training data.
Maintenance includes retraining models on fresh data, updating features, and sometimes rebuilding models entirely when business requirements change. Clear triggers for retraining (whether based on performance thresholds or time intervals) help teams stay proactive rather than reactive.
{{custom-cta-2}}
5 core components of ML model management
ML model management has many components that can vary depending on the architecture, goals, and environment. However, most ML model management strategies offer organizations five main components:
Versioning and code versioning
If you've ever worked in Google Docs and needed to consult a previous version of your file, you have some experience with versioning. Similarly, versioning for ML model management is when programmers keep track of files, changes made to the model and software, when and by whom, and what the results are.
More specific than just tracking versioning for models or data, code versioning is for the code used to create the models. Helpful for collaboration. Essential for creating easily repeatable models. And the only way you'll know exactly where things went wrong (or right) when training a model.
Together, versioning and code versioning enable reproducibility. When a model behaves unexpectedly, teams can trace back through the version history to understand what changed. When a model performs well, teams can recreate the exact conditions that produced it.
A minimum viable versioning approach should capture these elements:
For large language model (LLM)-based systems, versioning extends to prompts as well. Prompt templates should be versioned like code, with performance tracked by prompt version.
A reproducibility checklist helps teams verify they can recreate any model. To reproduce a model, you need the exact training data snapshot, the code commit that produced it, the environment configuration (dependencies and versions), the hyperparameters used, and the random seeds if applicable. Without these elements documented together, reproducing a model months later becomes guesswork.
Experiment tracking and model registry
An experiment tracker does just that: keeps track of experiments. Saving and organizing models. Recording metadata. Tracking the input and output of each model. Programmers can understand how the model works and why it gives the results that it does.
Once a model is finalized, programmers can store models in a central repository called a model registry. This helps keep your models organized. Additionally, a model registry can serve as a data set for training future models.
The model registry becomes the single source of truth for production models. It should include model artifacts, metadata about training data and parameters, performance metrics, lineage information showing how the model was created, and deployment status. When teams need to know which model version is running in production or compare current performance to previous versions, the registry provides those answers.
A well-structured registry entry typically contains:
- Model ID and semantic version
- Pointer to the training data snapshot
- Code commit hash and environment digest
- Evaluation metrics from testing
- Intended use documentation
- Approval status and approver identity
- Current deployment targets
Model monitoring and observability
Also called model observation, model monitoring is when programmers monitor models to make sure that their accuracy score stays at an acceptable level. Sometimes, models can experience "serving skew," which happens when there is a large discrepancy between the data a model was trained on and the data it's receiving in production. If the data sets are too different, the model will no longer be as accurate as it was during training.
Effective monitoring goes beyond simple accuracy tracking. Teams should monitor across four distinct layers:
- Data and feature monitoring: Track input data distributions, feature values, and data quality metrics to catch upstream issues before they affect predictions
- Model behavior monitoring: Watch prediction distributions, confidence scores, and output patterns for unexpected shifts
- Performance monitoring: Measure accuracy, precision, recall, or other task-specific metrics when ground truth labels become available
- System monitoring: Track latency, throughput, error rates, and resource utilization to ensure operational health
Data drift occurs when the statistical properties of input data change over time. Concept drift happens when the relationship between inputs and outputs shifts, even if the input distribution stays stable. Both require different detection approaches and response strategies. Teams sometimes conflate these two types of drift, but the distinction matters. Data drift might be addressed by retraining on recent data. Concept drift often requires revisiting the model architecture or feature set entirely.
Setting up automated alerts helps teams respond quickly when problems arise. Rather than discovering degraded performance weeks later, alerts can trigger when accuracy drops below a threshold, when prediction distributions shift unexpectedly, or when error rates spike.
Model governance and compliance
As machine learning moves from experimentation to enterprise-scale deployment, governance becomes essential. Organizations need to demonstrate that their models are fair, explainable, and compliant with regulations. Without the guardrails that AI governance frameworks provide, models become black boxes that create risk rather than value.
Governance encompasses several practices that work together:
- Documenting model purpose, intended use cases, and known limitations
- Establishing approval workflows before models reach production
- Maintaining clear ownership and accountability for each model
- Conducting regular audits of model performance and fairness
- Creating processes for handling model failures or unexpected behavior
Risk tiering adds another dimension to governance. Not every model needs the same level of oversight. A recommendation engine suggesting blog posts carries different risk than a model approving loan applications. Higher-risk models warrant stricter controls: more rigorous testing, additional approval gates, more frequent monitoring, and faster incident response requirements.
Compliance requirements increasingly apply to automated decision-making. Here's how common frameworks map to model management controls:
Building audit trails and lineage
Audit trails record every significant action in the model lifecycle. Who trained the model? What data did they use? When was it deployed? Who approved it? These records matter for compliance, but they also support explainable AI goals and help teams debug issues and understand how models evolved over time.
Model lineage goes deeper, tracking the full provenance of a model from raw data through feature engineering, training, and deployment. When a model makes a questionable prediction, lineage lets teams trace back to understand what data influenced that decision. Regulated industries like finance and healthcare increasingly require this traceability.
A practical audit log should capture the action taken (trained, evaluated, approved, deployed, rolled back), the actor (person or system), the timestamp, the model version affected, and any relevant context (approval notes, performance metrics at decision time, reason for rollback).
Role-based access and security
Not everyone in an organization should have the same access to ML systems. Data scientists might need full access to training environments but limited access to production. People in the business might need to view model outputs but not modify model configurations. Role-based access controls enforce these boundaries.
A typical access control matrix looks like this:
Security considerations extend beyond access control. Training data often contains sensitive information that requires strong data security controls. Model artifacts themselves can be valuable intellectual property. And deployed models can be targets for adversarial attacks, including data poisoning during training, model extraction attempts, and adversarial inputs designed to cause misclassification.
Model management for large language model (LLM) systems
Large language models introduce new management primitives that traditional ML frameworks don't address. The lifecycle stages remain similar. The artifacts and practices? Entirely different.
Prompt versioning deserves special attention. Prompts function like code in LLM systems, and small changes can dramatically affect outputs. Teams should version prompts in a registry, track performance metrics by prompt version, and run evaluation suites before promoting prompt changes to production. Teams often treat prompt changes as trivial tweaks that don't need the same rigor as code changes. A single word swap can shift model behavior in unexpected ways.
RAG (retrieval-augmented generation) systems add another layer of complexity. The retrieval index, the document corpus, and the embedding model all need versioning and governance. When an LLM produces an incorrect answer, teams need to trace whether the issue originated in retrieval (wrong documents surfaced), the prompt (poor instructions), or the model itself.
Evaluation harnesses for LLMs typically include accuracy tests against known-good answers, safety tests for harmful content generation, bias tests across demographic groups, and latency/cost benchmarks.
How to implement ML model management
Implementing ML model management strategies can be overwhelming, but the good news is that you can start small. Here are some of the core best practices to get you started on managing your organization's machine learning.
Logging and documentation best practices
Like a captain's log contains all the details about a ship and its voyage, ML model management logging contains details about the ML models. Logging should capture details about how programmers made the model, like model configurations, training parameters and hyperparameters, batch size, sampling techniques, and learning rate. Programmers should log metadata about the models, such as metrics, loss, configurations, and images. It is also helpful to log performance data, including confusion matrix, classification reports, and Shapley Additive Explanations (SHAP) values, both during training and after deployment.
Logging too much creates storage and performance problems. Logging too little makes debugging impossible. Find the balance by logging at meaningful checkpoints: epoch boundaries, validation runs, and deployment events.
Training and developing models effectively
Training and developing ML models starts with employees. Consider having peer reviews of training scripts. To simplify things while developing models, programmers should keep metric objectives simple and remove any features that are not being used. As far as the models themselves, lean heavily into versioning. This will help team members understand why the outputs are the way they are while providing greater insight into how models can be improved and developed.
ML models are fun to play around with, but to accomplish their objective, they need to have proper training metrics around model accuracy, deployment speed, and feedback loop efficiency. Programmers need clear, measurable, concrete objectives to make sure the models are working consistently. The metrics will vary depending on the task of the model, but some common ones are mean squared error (MSE), root mean squared error (RMSE), R², confusion matrix, F1 score, and gain and lift charts.
Deployment patterns and automation
Automate model deployment to reduce manual errors and shorten release cycles. You can also automate rollbacks for production models to increase efficiency. Even if you've done extensive testing, continue monitoring models after deployment to help you avoid plateaus and keep improving.
Deployment patterns vary based on use case requirements:
- Batch deployment: Works well for scenarios like nightly recommendation updates or weekly risk scoring, where predictions can be computed in advance and latency is not critical
- Real-time deployment: Suits applications requiring immediate responses, like fraud detection or dynamic pricing, where predictions must return in milliseconds
- Shadow deployment: Runs new models alongside existing ones without affecting people, allowing teams to compare performance before switching over
- Canary deployment: Routes a small percentage of traffic to the new model, gradually increasing if performance holds
Continuous integration and continuous delivery (CI/CD) practices from software development translate well to ML. Automated testing can validate model performance before deployment. Staged rollouts can limit exposure to potential issues. And automated rollback can quickly revert to previous versions when problems arise. Building these automation capabilities takes investment upfront but pays off as the number of models in production grows.
If you do not want to have to train every model yourself from scratch, you need to monitor and optimize model training. Models can train from other models, which makes the training process more efficient. You can optimize model training by experimenting with different model architectures, optimizers, loss functions, parameters, hyperparameters, and data.
Operational metrics and service level objectives (SLOs)
Treating ML systems like production software means defining service level objectives (SLOs) that set clear expectations for performance. These metrics vary by deployment pattern:
Quality thresholds depend on the use case, but common examples include accuracy above 95 percent, precision above 90 percent, and false positive rate below 5 percent.
Choosing ML model management tools
Selecting the right tools for model management depends on your organization's specific needs, existing infrastructure, and team capabilities. Rather than chasing the latest platform, focus on tools that address your actual pain points.
The build vs buy decision often comes down to team expertise and timeline. Open-source tools like MLflow and DVC offer flexibility and avoid vendor lock-in but require more engineering effort to operationalize. Cloud-native platforms from AWS, Azure, and Google provide integrated capabilities but can create dependency on a single provider. Commercial platforms offer governance and enterprise features but add licensing costs.
Key criteria to evaluate include:
- Integration with your existing data stack and cloud environment
- Support for your team's preferred languages and frameworks
- Collaboration features that match how your team works
- Governance capabilities that meet your compliance requirements
- Scalability to handle your expected model volume and complexity
- Total cost of ownership including licensing, infrastructure, and maintenance
A practical decision framework: if you need rapid experimentation and have ML expertise, start with open-source tools. If you need enterprise governance and have limited ML engineering capacity, evaluate cloud or commercial platforms. If you're already invested in a cloud provider, their native ML services often provide the smoothest integration path.
Look for platforms that make data AI-ready without requiring you to rebuild your infrastructure.
How Domo supports ML model management
Domo is here to support you as you invest in machine learning. The platform makes data AI-ready by connecting to your existing data sources and preparing that data for ML workflows. Rather than replacing your cloud data platform or preferred inference models, Domo works on top of them to accelerate time to value.
Governance runs throughout the Domo platform, ensuring that models built on your data maintain the access controls, audit trails, and compliance documentation that enterprise organizations require. When data scientists build models, they work with governed data. When those models deploy, they inherit the same governance framework.
Automated machine learning (AutoML) capabilities, built in partnership with Amazon Web Services, let teams train and tune ML models on governed data with human review and clear controls. With just a few clicks, AutoML prepares governed data for machine learning and launches training jobs that people can review before selecting a model for production.
Domo's AI agents can take model outputs and turn them into action on governed data, with human-in-the-loop review and controls that keep people in charge of objectives and constraints, distributing insights into the workflows people already use. Learn more about Domo's AI capabilities at ai.domo.com.




