Introduction to Machine Learning
1 Introduction
Machine learning (ML) is one of the fastest growing fields within Computer Science. For decades, building machines that could learn “on their own” was the long-standing dream of a small community of researchers. Today, ML systems are used in many domains, including medical diagnosis, translation, and game-playing, but also in controversial technologies like surveillance and automated weapon systems. So, while ML has produced wide-ranging benefits, the intentional and unintentional harms from its use and misuse cannot be ignored. The discussion of a “superintelligence” is also no longer limited to science fiction writers, as ML researchers consider safety implications of such systems before reaching key capability limits.
In past offerings of introductory ML courses, we have found that learners like you are taking the course for a variety of reasons. Many are interested in applying ML techniques in their field of interest. Others are hoping to gain fundamental skills to enable further study in the area. Still others are here with an open mind, with little prior expectation for the course but a lot of curiosity. You might even be here because of the hype you have seen in the news. We have also found that learners’ interests change: even if you are here to learn how to apply ML, you might find the mathematics and theory behind ML models fascinating—or vice versa. It’s good to keep an open mind and allow your goals and interests to change and grow with you.
These diverse goals, along with the constant development of the field, make it difficult to determine exactly what to teach in this course. Every day, new methods, models, and approaches are developed and presented. No doubt you have heard of Large Language Models (LLMs), Generative Artificial Intelligence (GenAI), and their many uses and implications. When you picture “AI” and “ML”, maybe a chat bot is what you have in mind, but this is just a tiny fragment of all that makes up AI and ML.
Instead of chasing the hype, our aim in this course is to learn the core, fundamental ideas of ML. These ideas have not changed much in the past few decades and yet underpin modern ML systems, including LLMs. So, although we will not focus on specific LLM architectures in this course, we will give you the mathematical and conceptual grounding that you will need if you plan to study them in greater depth.
In addition, our aim is for you to understand the core ideas that go beyond any single model that you see in the table of contents. We will return to these core ideas throughout the course; they are outlined below.
But first, it is important to cover a few key definitions. Terms like AI and ML are sometimes used interchangeably, even though they refer to different ideas. The next section will introduce these definitions.
2 Definitions
The term “Artificial Intelligence” (AI) was coined in 1956 by John McCarthy for the Dartmouth Summer Research Project on Artificial Intelligence. McCarthy proposed that “every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it.” The proposal for the project further states, “we think that a significant advance can be made in one or more of these problems if a carefully selected group of scientists work on it together for a summer.” While the problem turned out to be more complicated than that, the term “artificial intelligence” was coined.
Since then, several definitions of AI have been proposed. Here is one such definition, adapted from the textbook Artificial Intelligence: A Modern Approach by Russell & Norvig:
Definition: Artificial Intelligence (AI) is the study of agents that receive percepts from the environment and perform actions. The goal is to design agents that “behave rationally”, i.e., choosing the “best” action or making the “best” prediction, though “best” can mean different things in different contexts.
This is a broad definition for AI. “Agents that behave rationally” can mean different things in, say, playing chess, navigating a car, or predicting the presence of cancer with medical imaging. This definition also does not pose limits to any strategies used to build such agents. As such, what “AI” means has evolved to contain ideas from programming languages, algorithms, statistics, and other areas. You might be surprised to find that the textbook “Paradigms of Artificial Intelligence Programming” by Peter Norvig is actually a book about building (what are now considered rudimentary) AI agents using Common Lisp!
By contrast, ML defines a set of techniques that can be used to solve AI problems.
Definition: Machine Learning (ML) is a subset of artificial intelligence that enables systems to improve their performance on a task through experience. Machine learning algorithms allow computers to learn from and make decisions based on data, without being explicitly programmed for each specific task.
Rather than specifying the behaviour of an AI agent by hand, one can think of ML as automatically learning such a program from example data. That is, instead of writing a program (in, say, Python), we instead provide input-output examples, and use statistical techniques to identify patterns in these input-outputs.
Compared to programming, this is a fundamentally different way of interacting with computers and performing computing. As such, specifying computation using ML requires a different way of thinking. The kind of mistakes that ML systems can make are also different from programs specified using code. Thus, debugging strategies are vastly different in ML systems, along with other considerations.
There are many techniques within ML. We will study several of them in this course. There is one technique that you may have already heard of, related to the use of neural networks.
Definition: Deep Learning is a machine learning technique that uses multilayered neural networks to model complex patterns in data.
While deep learning is one successful (and interesting!) technique, it is not the only technique that we will study. In fact, deep learning techniques often build upon ideas and inspirations from other ML techniques. Understanding the foundations of ML, including simpler methods like nearest-neighbours and decision trees, provides crucial context for understanding how and why deep learning works.
Finally, we should mention Large Language Models, which have received significant attention in recent years.
Definition: Large Language Models (LLMs) are a class of deep learning models trained on vast amounts of text data to understand and generate human-like language.
LLMs represent a specific application of deep learning to natural language processing, demonstrating the power of these techniques when applied at scale. However, like deep learning more broadly, LLMs are built upon fundamental ML principles that we will explore throughout this course.
3 Fundamental Ideas in Machine Learning
As alluded to earlier, there are core ideas that recur across many ML models and methods discussed in this course. Although new models and methods are emerging constantly, these core ideas have stayed relatively constant. We consider these five ideas below to be “threshold concepts”: transformative, perspective-shifting ideas that characterize how ML practitioners approach problems. These ideas can be troublesome to learn and understand. It can take time to fully absorb them and their implications, and that is totally okay! We hope that you come back to this page and reflect on these ideas throughout your studies.
Idea #1. Learning Is Optimization
ML models “learn” a function that predicts a quantity that we care about. The term “learning” might evoke imagery of a complex psychological process that is indistinguishable from magic. In reality, in ML, we almost always turn learning problems into optimization problems. We will define what it means for a model to “perform well” via what we will call an objective function (or loss function), specify the set of possible models to choose from (e.g., models with different parameters or settings), and find the choice that “performs best” according to this objective. This approach turns a cognitive sounding concept like “learning” into a purely mathematical, mechanical process. Once understood, this idea of “learning as optimization” will demystify ML.
The focus on optimization is why mathematics is so important in ML. Finding the model that “performs best” requires using an optimization algorithm of some sort: perhaps a method based on derivatives and gradients like you have seen in a calculus course. Sometimes, a more naive exhaustive search could be used. Different learning algorithms will have different model settings and optimization algorithms associated with them.
As a consequence of treating learning as optimization, ML models are often highly modular. We can mix-and-match different combinations of mathematical representations (e.g. linear regression, as described in Chapter 3), objective functions, and optimization algorithms to produce different models. It may be that not all combinations perform equally well, but this modularity gives us many possible options to select from. This modularity also shapes how ML practitioners communicate with one another. Many ML papers introduce new models by defining a novel objective function, or a new optimization method for a specific problem. As a result, thinking about learning as optimization is a unified way of thinking about various ML problems, including new methods that are yet to be discovered.
Idea #2. ML Requires Balancing Tradeoffs in Sources of Error
Contrary to popular belief, ML models are not perfect; they can and will make mistakes. More interestingly, a more complex model will not necessarily make fewer mistakes! The combination of data choice (i.e., How much data do you have? What features are collected? How is the real-world problem framed as a learning problem), loss function (i.e., what does it mean for a model to “perform well”) and model choice all define the kinds of mistakes that models make.
What we mean by “kinds of mistakes” is that errors can be decomposed into three sources. For example, suppose that we are building an ML model that predicts a person’s height from their shoe size. Then, there could be error due to:
- the model being too simple for the given data set (e.g., imagine a very simple model that always predicts the same height: the average height across the population),
- the model being too complex and capturing patterns that happen to be in the data set but do not generalize (e.g., imagine that Tom happens to have a freakishly large shoe size, but average height; a complex enough model might capture the pattern that very large shoe size means average height),
- the variations in the problem itself that models cannot do anything about (e.g., because two people with the same shoe size can nevertheless have different heights, so the model cannot always predict the correct height for everyone).
Informally, the first two types of mistakes are related to the idea of underfitting/overfitting, which will be discussed throughout the text. Many models will have settings that tune the complexity of the model, so there is a tradeoff between the first two types of error, often called the bias-variance tradeoff, formalized in Chapter 6.
The third kind of error is tied to our problem formulation. If our model had more data features to work with (e.g. weight, in addition to shoe size), could it then predict the correct height for more people? This line of thinking might lead you to believe that it is always better for our data to have more features, but an overly rich data representation can actually result in a model that is too complex for the amount of training data we have. Both too few and too many data features highlight a fundamental aspect of ML: poor choice of input representation will lead to poor models.
In short, many limitations of ML systems cannot be solved through better (or more complex) algorithms. There are tradeoffs involved in designing ML systems, and a good system involves understanding the data, along with the model and choice of optimization criterion.
Idea #3. Model Evaluation Is Empirical
A related idea in ML is the understanding that no single algorithm is universally “best”. Formally, this is described in the “No Free Lunch” theorem (which we do not cover in this text). The implication is that the “best” ML approach is context dependent. Thus empirical evaluation using observed data is extremely important for ML practitioners.
What this means is that questions of the form “What choice of BLANK is best?” are often answered with “We need to try various options and see what works.” This is not satisfying, but it is a reality. In statistical learning, decisions are often made by making assumptions about the data distribution, and considering asymptotic or expected behaviours of various choices. However, in practical application, data is seldom independent and identically distributed, and rarely normally distributed. Asymptotic guarantees offer little comfort in safety critical systems. Empirical testing on realistic data is necessary.
We will also see that empirical testing is not straightforward. What exactly should we consider to be a “good” model? As we saw earlier, learning problems are turned into optimization problems, and typical optimization criteria tend to value models that produce “accurate” predictions across a data set. But what does “accurate” mean? Across what data set? For whom? For example, a model that predicts cancer risks may have equal “accuracy” across a dataset of men and women, but might over-estimate men’s cancer risks and under-estimate those of women. A facial recognition system may perform well across the data set that it is tested on, but have significantly higher error rates for darker-skinned individuals. A resume screening system may perform well on historical data, but only because it has encoded the same pattern of historical bias against under-represented groups. These examples are not hypothetical—they have all happened.
There are also additional empirical considerations beyond mere “accuracy”. Should our models be efficient? (And what does that mean?) Should our models be interpretable or explainable, so that we can understand why a model made the choices that it did? Should they be robust? Fair? Empirical evaluation also involves understanding how, why and for whom the model works, and what procedures need to be in place if models fail.
Idea #4. ML Describes Geometric Processes
The previous three ideas are more “applied” in nature. The next two ideas are more “theoretical”, and are important to understanding the mathematical foundations behind ML.
As mentioned earlier, our ML algorithms “learn” a function that predicts a quantity that we care about. They do so by leveraging patterns in a data set. A data set, as the name implies, is a set of data points. For example, we might curate a data set of the width and height of oak and maple leaves. Each data point within our data set has two measurements (the width and the height of some leaf) and therefore can map neatly into \(\mathbb{R}^2\). In more complex domains, a data point can have more properties (features) and the resulting data set can be treated as points in \(\mathbb{R}^D\). For example, images, music, and text can be treated as collections of points in \(\mathbb{R}^D\).
By representing data as points in a high-dimensional space, we can use geometric concepts like distance, similarity, and transformations. For example, two data points (leaves) that are “close” in this space (\(\mathbb{R}^2\)) are similar in some meaningful way, and two that are “far apart” are dissimilar. It matters less now what measures are being used. In fact, rotating the axes would not affect distances in this space, even though the numerical values of the features would change. This is the core idea behind distributed representations, including word and other embeddings used in language models. The measures along each axis become unimportant (and may not even be easily interpretable), but where the distances and relationships between data points encode key information.
The geometric perspective also helps us think of models not only as mappings from an input data point to an output prediction, but as transformations on entire data sets that manipulate the representations. Thus, thinking of data sets as geometric objects means thinking of models as geometric processes that act on a data set. This geometric viewpoint unifies seemingly different algorithms and provides intuition for how they work and when they might fail.
But models themselves can be thought of as geometric objects: certain models are more similar to each other than others. When conceived in this manner, the loss function and optimization processes can be thought of as geometric processes! When we optimize a model, we are essentially navigating through a high-dimensional space of possible models. This geometric view helps us understand and analyze optimizers, where they might succeed or fail, and where they might be efficient or inefficient.
To summarize, data, model, and loss can all be considered geometric objects and processes. This unified geometric perspective helps us see how all these components interact: the geometry of the data space influences what predictors can learn, the geometry of the model space determines what predictors are possible, and the geometry of the loss landscape determines how we can find good predictors.
This idea again highlights the importance of feature engineering and representation learning: the geometric structure we impose on data through our choice of features determines what kinds of patterns are easy or hard to discover.
Idea #5. ML Demands a Probabilistic Lens
In addition to data (and models, etc.) being geometric processes, a key concept in ML is understanding that deterministic reasoning about these processes is insufficient.
Real-world data is inherently uncertain and noisy, and can be thought of as sampling from a probability distribution. Thus, the data set that we have collected to train our models contains noise as well as the underlying generalizable pattern. As mentioned earlier, this uncertainty means we cannot expect our models to make perfect, deterministic predictions.
But paradoxically, understanding uncertainty means we can leverage it to build better models! Once we understand that data comes from probabilistic processes, we can design algorithms that explicitly account for this uncertainty. For example, one common strategy used in training models, called data augmentation, essentially adds our own noise to the collected data set. This mimics the natural variability we expect in real-world data and helps models become more robust to this variability. Understanding probabilistic processes helps us not only recognize sources of error, but also design better training procedures that account for uncertainty.
Many learning algorithms and loss functions themselves have probabilistic interpretations. Some models make these probabilistic interpretations and assumptions explicit, whereas others hide core assumptions behind seemingly deterministic procedures (see Chapter 9 on k-means clustering). Advanced ML models are often reasoned about through their probabilistic interpretation of data and its generating processes, in order to build better models.
This probabilistic perspective also appears in optimization procedures. Many optimization algorithms used in ML are stochastic in that they incorporate randomness in their search process. Again, this randomness is a source of strength rather than a weakness: introducing noise in certain optimization processes can actually help the algorithm find better solutions, making the optimization more robust to the specific characteristics of the training data. This randomness isn’t just a computational convenience; it reflects the probabilistic nature of the data itself.
The geometric and probabilistic perspectives complement each other: geometry helps us understand the structure and relationships in data, while probability helps us reason about uncertainty and variability. Together, they form the mathematical foundation of ML. Understanding both perspectives helps us design better models, choose appropriate algorithms, and reason about when and why our models succeed or fail.
4 Summary
We hope these notes help you find ML interesting and accessible. With ML becoming so impactful, we hope to support you in whatever goals you find meaningful: whether it is becoming critical users of ML technologies, thoughtful builders of ML tools, researchers of novel ML methods —or paths we haven’t imagined yet. We look forward to learning alongside you.