Join the AI + Data Tour for hands-on training, real customer stories, and time with Domo product experts near you.
Machine Learning Basics: A Practical Guide for 2026

Machine learning powers everything from spam filters to sales forecasts. Yet the path from raw data to working model remains murky for many teams. This guide breaks down the fundamentals: supervised, unsupervised, and reinforcement learning approaches; the end-to-end pipeline that turns data into predictions; evaluation metrics that actually matter; and the mistakes that derail projects before they reach production. By the end, you'll understand not just what machine learning is, but how to start applying it to problems that matter to your organization.
Key takeaways
Here are the main points to keep in mind:
- Machine learning trains systems to improve at tasks through data patterns rather than explicit programming
- The three main types are supervised learning (labeled data), unsupervised learning (pattern discovery), and reinforcement learning (trial and error)
- The machine learning (ML) pipeline includes problem framing, data preparation, model training, evaluation, and deployment with monitoring
- Getting started requires understanding data preparation as the foundation, starting with simple baseline models before advancing
- Success depends on choosing the right evaluation metrics for your problem type and avoiding common pitfalls like data leakage
What is machine learning?
Machine learning is a branch of artificial intelligence that trains systems to improve their performance on tasks without being explicitly programmed for every scenario. Instead of following rigid instructions, machine learning algorithms learn from data, identify patterns, and make decisions based on what they discover.
The core idea is straightforward. Feed a system enough examples, and it starts recognizing what matters. Show it thousands of email samples labeled as spam or not spam, and it learns to spot the difference on its own. Give it years of sales data, and it can predict which deals are most likely to close.
This approach flips traditional software development on its head. Rather than a programmer anticipating every possible situation and writing rules to handle each one, machine learning lets the data teach the system what to do.
How machine learning differs from traditional programming
Traditional programming follows a simple formula: a programmer writes code that takes data as input and produces a specific output. Every rule, every exception, every edge case needs to be anticipated and coded manually.
Machine learning reverses this relationship. The system takes both data and the desired output, then creates its own program (the model) for new situations. The machine figures out the rules instead of a human spelling them out.
Why does this distinction matter for business applications? Traditional programming works well when rules are clear and don't change often. Machine learning shines when patterns are complex, data is abundant, and conditions shift regularly. As a result, machine learning helps increase the value of embedded analytics, speeds up insights for people, and reduces decision bias.
Machine learning vs artificial intelligence
While machine learning is a subset of artificial intelligence, the two terms describe different scopes. Artificial intelligence is the broader goal of enabling machines to think and make decisions as a human would. Machine learning is one method for achieving that goal, specifically by training systems to improve at tasks through exposure to data.
Think of AI as the destination and machine learning as one of the most effective routes to get there.
Machine learning vs deep learning
Deep learning is a subset of machine learning that uses multi-layered neural networks to process information. Designers modeled these networks on the human brain, with layers of interconnected nodes that progressively extract higher-level features from raw input.
Image and speech recognition made deep learning famous. It excels at finding complex patterns in large amounts of unstructured data. While traditional machine learning algorithms might require humans to identify which features matter, deep learning figures out the relevant features on its own.
Machine learning vs data science
Data science is a broader discipline that encompasses the entire process of extracting insights from data, including statistics, data visualization, domain expertise, and communication. Machine learning is one tool within the data science toolkit.
A data scientist might spend weeks cleaning data, exploring patterns, and building visualizations before deciding whether machine learning is even the right approach. When ML is appropriate, the data scientist applies it as part of a larger analytical workflow. Data science asks what a team can learn from the data, while machine learning asks whether a team can train a system to make predictions or decisions from the data.
When to use each approach
Choosing between traditional programming, machine learning, deep learning, and broader data science depends on your problem characteristics and available resources. The following decision rules help match approaches to situations:
- Use traditional programming when rules are clear, don't change often, and can be explicitly defined (calculating taxes, validating form inputs, routing logic)
- Use machine learning when patterns are complex, data is abundant, and you need predictions on structured data (churn prediction, demand forecasting, fraud scoring)
- Use deep learning when working with unstructured data like images, audio, or text, and you have large datasets plus computational resources
- Use data science when you need to explore data, generate insights, and communicate findings (ML may or may not be part of the solution)
If you can write down the rules in an afternoon, traditional programming is simpler and more maintainable.
3 types of machine learning
Machine learning approaches fall into three main categories, each suited to different problems and data situations.
Supervised learning
Supervised learning is the most common type and powers most production ML systems today. The name comes from the learning process: the algorithm trains on labeled data where the correct answer is already known, like a student learning from a teacher who provides the right answers.
This type of learning handles two main tasks. Regression predicts numerical values, using age and income history to forecast future earnings, for instance. Classification assigns categories, like determining whether a customer will make a specific purchase based on their browsing behavior. Treating classification problems as regression (or vice versa) leads to models that optimize for the wrong objective entirely.
Common supervised learning algorithms include:
- Linear regression
- Logistic regression
- Decision trees
- Random forest
- Gradient boosting
- Artificial neural networks
The Domo Data Science suite provides tools for building and deploying supervised learning models with governed workflows, human oversight, and human-in-the-loop review.
Unsupervised learning
Unsupervised learning works with unlabeled data, discovering patterns and relationships without being told what to look for. This approach proves valuable when you have data but don't know what questions to ask yet.
Clustering is the most common application. The algorithm groups similar items together. Customer segmentation is a classic example: feed the system purchase history and behavior data, and it identifies natural groupings you might not have noticed. K-Means is the most widely used clustering algorithm, though it requires you to specify the number of clusters upfront. Choosing this number arbitrarily rather than testing multiple values often produces meaningless segments.
Unsupervised learning also handles dimensionality reduction, which simplifies complex datasets while preserving important information.
Reinforcement learning
Reinforcement learning takes a different approach entirely. Instead of learning from examples, the system learns through trial and error, receiving rewards for desirable actions and penalties for mistakes.
This type of learning excels in situations where the optimal action depends on a sequence of decisions. Robotics, game playing, and autonomous vehicle navigation all rely heavily on reinforcement learning. The system explores different strategies, learns which ones lead to improved outcomes, and gradually improves its decision-making.
Choosing the right type for your problem
Selecting the appropriate ML type depends on your data and goals:
- Use supervised learning when you have labeled examples and want to predict outcomes for new data (customer churn, fraud detection, demand forecasting)
- Use unsupervised learning when you want to discover structure in data without predefined categories (customer segmentation, anomaly detection, topic modeling)
- Use reinforcement learning when optimal actions depend on sequential decisions and you can define a reward signal (robotics, game AI, resource optimization)
If you have historical data with known outcomes, start with supervised learning. Reinforcement learning typically requires specialized infrastructure and is less common in standard business applications.
How machine learning works
Understanding the machine learning workflow helps demystify what happens between raw data and useful predictions.
The machine learning pipeline
Every machine learning project moves through these stages:
- Define the problem: Clarify what you're trying to predict or discover, and determine whether machine learning is the right approach
- Collect data: Gather relevant data from available sources, ensuring you have enough examples to train a reliable model
- Prepare data: Clean, transform, and organize the data so algorithms can process it effectively
- Select and train a model: Choose an appropriate algorithm and let it learn patterns from your prepared data
- Evaluate performance: Test the model on data it hasn't seen to measure how well it generalizes
- Deploy and monitor: Put the model into production, applying model management practices to track its performance over time
Data preparation typically consumes the most time in any ML project. That part often gets too little attention. Domo's extract, transform, and load (ETL) tools help integrate, clean, and transform data, addressing one of the most challenging parts of the data-to-analysis process.
Key elements of every ML algorithm
Three elements define how any machine learning algorithm operates:
- Representation: What the model looks like and how knowledge is stored
- Evaluation: How good models are distinguished from poor ones and how performance is measured
- Optimization: The process for finding the best model parameters and improving results
Domo has created a Machine Learning playbook that walks through properly preparing data, running a model in a ready-made environment, and visualizing results. Since building and choosing a model can be time-consuming, automated machine learning (AutoML) handles much of this work automatically, including data preprocessing, model selection, and hyperparameter tuning.
Core concepts beginners need to understand
Several foundational concepts trip up newcomers to machine learning. Getting these right early prevents confusion and debugging headaches later.
Features and labels
Features are the input variables your model uses to make predictions. Labels are the outputs you're trying to predict. In a house price prediction model, features might include square footage, number of bedrooms, and neighborhood. The label is the sale price.
Think of features as the columns in a spreadsheet that describe each example, and labels as the column containing the answer you want to predict. The quality and relevance of your features often matter more than which algorithm you choose.
A simple example:
The first three columns are features. The last column is what the model learns to predict.
Parameters vs hyperparameters
Parameters are values the model learns during training. In a linear regression, the coefficients that multiply each feature are parameters. The model adjusts these automatically to minimize prediction errors.
Hyperparameters are settings you choose before training begins. Learning rate, number of trees in a random forest, or the number of layers in a neural network are hyperparameters. These control how the model learns rather than what it learns.
Here's the analogy: parameters are like the answers on a test that the student figures out. Hyperparameters are like the study conditions you set beforehand (how long to study, which textbook to use, how many practice problems to attempt).
Training, validation, and test data
Machine learning projects split data into three sets:
- Training data: The examples your model learns from (typically 60 to 80 percent of available data)
- Validation data: Used to tune hyperparameters and compare model variations during development (typically 10 to 20 percent)
- Test data: Held out completely until final evaluation to estimate how the model will perform on new data (typically 10 to 20 percent)
Never use test data during model development. If you peek at test results and adjust your model accordingly, you lose the ability to estimate how it will perform in production.
Generalization and the bias-variance tradeoff
Generalization is the ability of a model to perform well on data it hasn't seen before. A model that memorizes training examples but fails on new data has poor generalization.
The bias-variance tradeoff describes the tension between two types of errors. High bias means the model is too simple and misses important patterns (underfitting). High variance means the model is too sensitive to training data specifics and doesn't generalize (overfitting). Finding a model complex enough to capture real patterns but simple enough to avoid memorizing noise? That's the goal.
Think of it like studying for an exam. Memorizing every practice problem word-for-word (high variance) won't help if the exam has different questions. But only learning the chapter titles (high bias) means missing the details you need.
Common machine learning challenges
Several problems plague machine learning projects more than any others.
Overfitting
Overfitting happens when a model learns the training data too well, including its noise and random fluctuations. The model performs brilliantly on training data but fails when it encounters new examples. Signs of overfitting include:
- Near-perfect accuracy on training data but poor performance on test data
- Model complexity that far exceeds what the problem requires
- Predictions that seem to memorize specific examples rather than learn general patterns
Common fixes include using more training data, simplifying the model, applying regularization techniques, or using cross-validation to detect the problem early.
Underfitting
Underfitting is the opposite problem. The model is too simple to capture the underlying patterns. It performs poorly on both training and test data because it hasn't learned enough from the examples provided. Signs include:
- Consistently poor performance across all datasets
- Model predictions that miss obvious patterns humans can see
- High error rates that don't improve with more training
Fixes for underfitting include using a more complex model, adding more relevant features, reducing regularization, or training for more iterations.
Data leakage
Data leakage occurs when information from outside the training dataset influences the model, leading to overly optimistic performance estimates that don't hold up in production. Common sources:
- Using features that wouldn't be available at prediction time (like including tomorrow's stock price to predict today's)
- Preprocessing data before splitting into train/test sets (like normalizing using statistics from the entire dataset)
- Target leakage where the label directly or indirectly appears in the features
Leakage is particularly dangerous because models appear to work well during development but fail when deployed. Always ask: "Would I have access to this information at the moment I need to make a prediction?"
Common beginner mistakes and fixes
Beyond the major challenges, several specific mistakes trip up newcomers. Here's what to watch for and how to address each one:
- Not splitting data before preprocessing: Scaling or encoding using the full dataset leaks test information into training. Fix: Always split first, then fit transformers on training data only.
- Using accuracy on imbalanced data: A model predicting "not fraud" for everything achieves 99 percent accuracy when 99 percent of transactions are legitimate, while catching zero fraud. Fix: Use precision, recall, F1, or receiver operating characteristic area under the curve (ROC-AUC) instead.
- Forgetting to set random seeds: Results aren't reproducible across runs. Fix: Add randomstate=42 (or any fixed number) to traintestsplit and model initialization.
- Training on too little data: Models can't learn patterns from dozens of examples. Fix: Gather more data, use data augmentation, or try simpler models.
- Ignoring feature scaling: Algorithms like k-nearest neighbors (k-NN) and gradient descent perform poorly when features have vastly different scales. Fix: Standardize or normalize features before training.
Model evaluation essentials
Choosing the right evaluation metrics and avoiding common pitfalls determines whether your model actually solves the problem or just appears to.
Metrics for classification
Classification tasks predict categories, and different metrics capture different aspects of performance:
- Accuracy: Percentage of correct predictions. Useful when classes are balanced, misleading when they're not.
- Precision: Of all positive predictions, how many were actually positive? Important when false positives are costly.
- Recall: Of all actual positives, how many did the model catch? Important when false negatives are costly.
- F1 Score: Harmonic mean of precision and recall. Useful when you need to balance both.
- ROC-AUC: Measures the model's ability to distinguish between classes across all threshold settings.
For imbalanced datasets (like fraud detection where 99 percent of transactions are legitimate), accuracy is nearly useless. A model that predicts "not fraud" for everything achieves 99 percent accuracy while catching zero fraud cases.
The confusion matrix helps visualize these tradeoffs. It shows four quadrants: true positives (correctly predicted fraud), true negatives (correctly predicted legitimate), false positives (flagged legitimate as fraud, annoying customers), and false negatives (missed fraud, costly losses). Which errors matter more depends entirely on your business context.
Metrics for regression
Regression tasks predict numerical values, with these common metrics:
- Mean Absolute Error (MAE): Average absolute difference between predictions and actual values. Easy to interpret in original units.
- Mean Squared Error (MSE): Average squared difference. Penalizes large errors more heavily than MAE.
- Root Mean Squared Error (RMSE): Square root of MSE. Same units as the target variable.
- R-squared: Proportion of variance explained by the model. Ranges from 0 to 1, with higher being better.
If large errors are particularly problematic, MSE or RMSE makes sense. If you want a metric that's easy to explain to stakeholders, MAE often works well.
Cross-validation
Cross-validation provides more reliable performance estimates than a single train/test split. The most common approach, k-fold cross-validation, works as follows:
- Split data into k equal parts (typically five or 10)
- Train on k-1 parts, evaluate on the remaining part
- Repeat k times, each time holding out a different part
- Average the results across all k runs
This approach uses all data for both training and evaluation while maintaining separation. It's particularly valuable when data is limited.
Evaluation checklist
Before trusting your model's performance score, verify these conditions:
- Used a holdout test set that the model never saw during training or tuning
- Checked for data leakage by confirming all features would be available at prediction time
- Chose metrics appropriate for the task type and class balance
- Compared performance to a simple baseline (mean prediction, majority class, or simple rules)
- Validated that performance is consistent across cross-validation folds
Algorithm selection and baseline strategy
Dozens of algorithms available. Where do you start?
Start with a baseline
Before trying sophisticated algorithms, establish a baseline using the simplest reasonable approach. For classification, this might be predicting the most common class. For regression, predicting the mean value. This gives you a floor to beat.
Next, try a simple interpretable model. Linear regression for regression tasks, logistic regression for classification. These models train quickly, are easy to debug, and often perform surprisingly well. If a simple model solves your problem adequately, you're done.
Why baseline first? A simple model establishes minimum acceptable performance. If your complex model doesn't beat the baseline, something is wrong with your data, features, or problem framing.
When to use common algorithms
The following guidance helps match algorithms to problem characteristics:
- Linear/Logistic Regression: Start here. Works well with many features, handles large datasets efficiently, and produces interpretable results. Best when relationships are roughly linear.
- Decision Trees: Good for understanding feature importance and handling non-linear relationships. Prone to overfitting but easy to visualize and explain.
- Random Forest: Ensemble of decision trees that reduces overfitting. Strong general-purpose choice for tabular data. Less interpretable than single trees but more stable.
- Gradient Boosting (XGBoost, LightGBM): Often achieves top performance on structured data. More complex to tune but frequently wins competitions and production benchmarks.
- K-Nearest Neighbors: Simple and intuitive. Works well with smaller datasets but slows down significantly as data grows.
- Neural Networks: Powerful for unstructured data (images, text, audio) and very large datasets. Requires more data and compute than other approaches. Less interpretable.
For most business problems with structured data, start with logistic regression or random forest.
Algorithm selection heuristics
These rules of thumb help narrow choices:
- Small dataset (fewer than 1,000 samples): Try simpler models like linear regression, logistic regression, or k-NN
- Large structured dataset: Random forest or gradient boosting typically perform well
- Images, text, or audio: Consider deep learning if you have sufficient data and compute
- Need interpretability: Use linear models or decision trees
- Need maximum accuracy regardless of complexity: Try ensemble methods or neural networks
Machine learning applications across industries
Machine learning has moved from research labs into everyday business operations. The following table shows how different industries apply these techniques:
Business and enterprise applications
Machine learning helps businesses reach their desired outcomes more quickly. Organizations use ML to improve efficiencies and operations, perform preventative maintenance, adapt to changing market conditions, and use consumer data to increase sales and improve retention.
The most successful implementations focus on specific, measurable problems. Rather than pursuing ML for its own sake, effective teams identify where predictions would change decisions and work backward from there. You'll notice the organizations that struggle are usually the ones that start with an AI mandate instead of a clear problem to solve.
Consumer-facing applications
Streaming platforms and e-commerce websites harness machine learning to deliver personalized content and product suggestions tailored to people's preferences and behavior. These recommendation systems analyze viewing history, purchase patterns, and similar behavior across people to surface relevant options.
Computer vision technology powers applications like facial recognition, object detection, and image classification, playing a vital role in security systems and autonomous vehicles. Virtual assistants and transcription tools rely on machine learning to accurately convert spoken language into text.
How machine learning models improve over time
Machine learning models refine their predictions and accuracy by continuously learning from new data. The key mechanisms for improvement include:
- Retraining with fresh data: Teams retrain models as new data becomes available to improve performance
- Fine-tuning hyperparameters: Adjusting algorithm parameters enhances prediction accuracy
- Transfer learning: Teams adapt pre-trained models to new but related tasks to accelerate learning
- Eliminating biases: Continuous monitoring ensures models remain fair and accurate over time
Advancements in model architectures, such as transformers and neural networks, enable more efficient learning and improved handling of complex data. Improved computational power and scalable infrastructure also play a critical role in speeding up the training and optimization processes.
Getting started with machine learning
Breaking into machine learning doesn't require a PhD or years of specialized training. The field has become increasingly accessible.
Start with these foundational steps:
- Learn the basics of Python or R, the two most common languages for ML work
- Understand statistics fundamentals: mean, variance, probability distributions, and hypothesis testing
- Practice with structured datasets before tackling messy data
- Use AutoML tools to build initial models while learning what happens under the hood
- Focus on one problem type (classification or regression) until you understand it well
Data preparation deserves special attention. Most ML projects spend 60 to 80 percent of their time cleaning and transforming data. This ratio holds true whether you're a beginner or a seasoned practitioner. Building strong data wrangling skills pays dividends throughout your ML journey.
Platforms like Domo reduce the barrier to entry by handling infrastructure complexity and providing AutoML capabilities with governance controls, human oversight, and human-in-the-loop model review. This lets teams focus on understanding their data and interpreting results rather than managing technical infrastructure.
Self-learning roadmap and realistic expectations
Can you learn machine learning by yourself? Yes. But expect three to six months to reach basic competency, not days. A realistic roadmap looks like this:
- Month 1: Python fundamentals plus pandas and NumPy for data manipulation
- Month 2: Core ML concepts and hands-on practice with scikit-learn
- Month 3: First complete project from problem framing through evaluation
- Months 4-6: Specialization in an area that interests you (natural language processing, computer vision, time series)
What you can learn in one week: terminology, one algorithm (linear regression), and one end-to-end mini-project. What you cannot learn in one week: deep learning, production deployment, or advanced mathematical foundations. Use that first week to decide if you want to continue.
Prerequisites are lighter than most people assume. You need basic Python (variables, loops, functions), high school math (algebra and basic statistics), and five to 10 hours per week of study time. Start with Python, Jupyter notebooks, scikit-learn, and pandas. Do not worry about TensorFlow or cloud platforms until you have built several working models with simpler tools.
The future of machine learning
Machine learning continues to evolve rapidly. Improvements in unsupervised learning algorithms will contribute to more accurate analysis, which will inform clearer insights.
Natural language processing has advanced dramatically. Systems now understand context, nuance, and intent in ways that seemed impossible just a few years ago. Search systems can interpret different kinds of queries and provide more accurate answers, changing how people interact with information.
The rise of agentic AI represents the next frontier. Rather than models that simply respond to prompts, agentic systems can plan, execute multi-step tasks, and coordinate with other systems to achieve goals. This shift requires careful attention to governance and human oversight, ensuring that automated systems operate within defined boundaries. The complexity here is real, and it is not going away anytime soon.
Marketing teams are increasingly adopting artificial intelligence and machine learning to improve personalization strategies and customer engagement. When combined with strong data foundations, machine learning can provide insights that propel a company forward.
The organizations seeing the most value from ML are those treating it as a capability to build rather than a project to complete.




