Logistic Regression

Learning Objectives

After reading this page, you should be able to:

  1. Derive weights to perfectly classify a simple logical function using binary linear classification.
  2. Explain the problems with using the 0-1 loss, squared loss, and sigmoid activation function + squared loss for binary linear classification.
  3. Define the logistic regression model (linear model, logistic activation function, and cross-entropy loss).
  4. Apply logistic regression to solve a binary classification problem.

1 Introduction

Thus far, we have learned how to solve a regression problem using a linear model. What about a classification problem, that is, when the target variable is categorical? In this chapter, we will extend linear regression and apply the resulting model, called logistic regression, to classification problems. Despite its name, logistic regression is a classification model, not a regression model. The name is historical: the model fits a linear function to a continuous quantity, which we then convert into a class prediction. We will start with a binary classification problem, where the target variable takes one of two possible values. The leaf prediction problem introduced earlier is an example of a binary classification problem.

This page will take us through the journey of designing the logistic regression model, a widely used model for binary classification. Our journey will follow a principled approach. We begin by considering simpler alternatives that will ultimately fail. By analyzing why these approaches fail, we will gain insight into the design choices of the logistic regression model. This journey will illustrate how the different components in the modular approach, such as the predictor, the loss, and the optimizer, interact with one another in interesting ways. More importantly, this process illustrates a key lesson in machine learning: model design involves carefully considering the properties we need from the model architecture and the loss function.

2 Binary Classification Problem

We will use the leaf prediction dataset as an example of a binary classification problem. Recall that each leaf is described by two measurements: width and height. Our task is to predict whether each leaf is an Oak (Oak) or Maple (Maple). The training data consists of seven labelled examples:

Table 1: Training data for the leaf prediction problem.
# Leaf Width Leaf Height Leaf Type
1 7.0 cm 12.0 cm Oak
2 9.0 cm 6.0 cm Oak
3 10.0 cm 18.0 cm Oak
4 5.0 cm 10.0 cm Oak
5 13.0 cm 11.0 cm Maple
6 11.0 cm 9.0 cm Maple
7 9.0 cm 14.0 cm Maple

Formally, we can denote the dataset using the same notation as before, where each target \(t^{(i)}\) is categorical.

\[\begin{align*} \Big\{ \left( \mathbf{x}^{(1)}, t^{(1)} \right), \left( \mathbf{x}^{(2)}, t^{(2)} \right), \dots, \left( \mathbf{x}^{(N)}, t^{(N)} \right) \Big\} \end{align*}\]

Since we focus on binary classification, the target \(t^{(i)}\) can take one of two values, \[\begin{align*} t^{(i)} \in \{0, 1\} \end{align*}\] Following convention, we refer to \(t^{(i)} = 1\) as a positive example and \(t^{(i)} = 0\) as a negative example. In the leaf dataset, we might encode Oak as \(t^{(i)} = 0\) (negative) and Maple as \(t^{(i)} = 1\) (positive), though this assignment is arbitrary.

3 Geometric Intuition for Classification Models

Let’s modify the linear regression model slightly so that it can be used for binary classification. The main goal of this section is to build some geometric intuition about decision boundaries, the data space, and the weight space.

Recall that the linear regression model computes predictions as a weighted combination of input features. Formally, the prediction \(z\) is given by
\[\begin{align*} z = \left( \sum_{j=1}^D w_j x_j \right) + b = \mathbf{w}^{\top} \mathbf{x} + b. \end{align*}\] So far, the prediction \(z\) is a real number. For binary classification, we need to generate a binary prediction \(y \in \{0, 1\}\). One easy approach to convert a real number to a binary number is to apply a threshold \(r\). \[\begin{align*} y = \begin{cases} 1 & \text{if } z \geq r \\ 0 & \text{if } z < r \end{cases} \end{align*}\] Combining these two steps gives us our first binary classification model: \[\begin{align*} z = \mathbf{w}^{\top} \mathbf{x} + b \qquad\qquad y = \begin{cases} 1 & \text{if } z \geq r \\ 0 & \text{if } z < r \end{cases} \end{align*}\] where

  • \(\mathbf{w} \in \mathbb{R}^D\) is the weight vector with \(D\) features,
  • \(\mathbf{x} \in \mathbb{R}^D\) is the input vector with \(D\) features,
  • \(b\) is the bias term,
  • \(r\) is the threshold,
  • \(z\) is an intermediate real-valued quantity that we compute, and
  • \(y\) is the binary prediction.

Question: We can simplify this formulation further. Show that the original model is equivalent to the following model: \[\begin{align*} z = \mathbf{w}^{\top} \mathbf{x} \qquad\qquad y = \begin{cases} 1 & \text{if } z \geq 0 \\ 0 & \text{if } z < 0 \end{cases} \end{align*}\]

Answer:

First, we can set the threshold \(r\) to \(0\). Any condition \(\mathbf{w}^{\top} \mathbf{x} + b \geq r\) can be rewritten by moving \(r\) into the bias: \[\begin{align*} \mathbf{w}^{\top} \mathbf{x} + b \geq r \Rightarrow \mathbf{w}^{\top} \mathbf{x} + (b - r) \geq 0 \Rightarrow \mathbf{w}^{\top} \mathbf{x} + b' \geq 0 \end{align*}\]

Intuitively, the original model has two offset parameters, the bias \(b\) and the threshold \(r\), both affecting the prediction. However, only their difference \(b - r\) affects the decision rule, so we replace them with a single bias \(b' = b - r\) and set the threshold to \(0\).

The resulting model becomes: \[\begin{align*} z = \mathbf{w}^{\top} \mathbf{x} + b \qquad\qquad y = \begin{cases} 1 & \text{if } z \geq 0 \\ 0 & \text{if } z < 0 \end{cases} \end{align*}\] Second, we can absorb the bias term into the weight vector by adding a dummy feature \(x_0 = 1\) to the input vector. This step is the same as the one we performed for linear regression. Both vectors now have \(D+1\) entries, so \(\mathbf{x}, \mathbf{w} \in \mathbb{R}^{D+1}\), where the extra entries are the dummy feature and the bias. The model becomes: \[\begin{align*} z = \mathbf{w}^{\top} \mathbf{x} \qquad\qquad y = \begin{cases} 1 & \text{if } z \geq 0 \\ 0 & \text{if } z < 0 \end{cases} \end{align*}\]

3.1 The Decision Boundary

Let’s consider the decision boundary for this model. Recall that the decision boundary is the set of points where the model’s prediction would change if the input changes a little bit.

Question: Determine the equation of the decision boundary for the simplified model below. Briefly explain the shape of the decision boundary.

\[\begin{align*} z = \mathbf{w}^{\top} \mathbf{x} \qquad\qquad y = \begin{cases} 1 & \text{if } z \geq 0 \\ 0 & \text{if } z < 0 \end{cases} \end{align*}\]

Answer:

The prediction changes when \(z = 0\). The decision boundary is therefore the set \[\begin{align*} \{\mathbf{x} : \mathbf{w}^{\top} \mathbf{x} = 0\}. \end{align*}\] This set is a hyperplane. In 1D, it is a point. In 2D, it is a line. In 3D, it is a plane. In higher dimensions, it is a hyperplane.

To make this concrete, we can visualize a decision boundary for such a linear classification model trained on the leaf dataset. Recall that we represent each leaf using two features: \(x_1\) (leaf width in cm) and \(x_2\) (leaf height in cm), along with the constant feature \(x_0 = 1\) for the bias term.

Consider a linear classification model with the weights below. What does the decision boundary for this model look like?

\[w_0=-150 \text{ (bias)}, w_1=5 \text{ (weight for width)}, w_2=10 \text{ (weight for height)}\]

With these weights, the decision rule is: \[\begin{align*} z &= w_0 + w_1 x_1 + w_2 x_2 = -150 + 5 x_1 + 10 x_2 \\ y &= \begin{cases} 1 \text{ (Maple)} & \text{if } z \geq 0 \\ 0 \text{ (Oak)} & \text{if } z < 0 \end{cases} \end{align*}\]

The decision boundary occurs where \(z = 0\), which means \(-150 + 5 x_1 + 10 x_2 = 0\), or equivalently, \(x_2 = 15 - 0.5 x_1\). This is a line in the \((x_1, x_2)\) feature space that separates leaves we classify as Oak from those we classify as Maple.

As the choice of the weights \(\mathbf{w}\) changes, so does the decision boundary. As in linear regression, it can be useful to visualize the weight space to obtain a geometric understanding of how the decision boundaries change with the weights. The following visualization shows the decision boundary in the data space on the left and in the weight space on the right while fixing \(w_0 = -150\).

Decision boundary in data space and weight space

Figure 1: Interactive Click in the weight space to select different weights and see how the decision boundary changes in data space. Shown for linear classification with \(w_0 = -150\) fixed. The data space (left) shows the decision boundary for the selected weights, initially set to \(\begin{bmatrix} w_1 & w_2 \end{bmatrix}^\top = \begin{bmatrix} 5 & 10 \end{bmatrix}^\top\), which corresponds to the decision boundary \(x_2 = 15 - 0.5 x_1\).

Just like in linear regression, the data-space visualization gives a geometric way to judge a set of weights. A good set of weights will correctly classify many training data points, ideally all of them. The next question is what this requirement looks like in weight space.

3.2 Constraints in the Weight Space

We can answer this question by examining how individual data points constrain the weights. Each training example imposes a constraint on the weight vector. For our leaf dataset, consider leaf #1: \[\text{features: } \begin{bmatrix} x_0^{(1)} \\ x_1^{(1)} \\ x_2^{(1)} \end{bmatrix} = \begin{bmatrix} 1 \\ 7 \\ 12 \end{bmatrix} \quad \text{label: } t^{(1)} = 0 \text{ (Oak)}.\] For the model to correctly classify this leaf, we need the weights to satisfy the following inequality: \[\begin{align*} z = w_0 + w_1 x_1^{(1)} + w_2 x_2^{(1)} = w_0 + w_1 \cdot 7 + w_2 \cdot 12 < 0 \end{align*}\]

This inequality defines a half-space in the three-dimensional weight space \((w_0, w_1, w_2)\). Our chosen weights \(\begin{bmatrix} w_0 & w_1 & w_2 \end{bmatrix}^\top = \begin{bmatrix} -150 & 5 & 10 \end{bmatrix}^\top\) fail to satisfy this constraint, as shown below: \[-150 + 5 \cdot 7 + 10 \cdot 12 = -150 + 35 + 120 = 5 \not< 0\] This means that our example weights would misclassify leaf #1.

To visualize these constraints more clearly, we can fix the bias \(w_0 = -150\) and examine how the data constrains the remaining weights \(\begin{bmatrix} w_1 & w_2 \end{bmatrix}^\top\) in a two-dimensional weight space. Each training example then defines a half-space in this weight space. For example, the constraint for the oak leaf #1 is shown below: \[\begin{align*} \begin{bmatrix} x_0^{(1)} \\ x_1^{(1)} \\ x_2^{(1)} \end{bmatrix} = \begin{bmatrix} 1 \\ 7 \\ 12 \end{bmatrix} \qquad \text{ requires } \qquad 7w_1 + 12w_2 < 150 \end{align*}\]

The constraint for the maple leaf #5 is shown below: \[\begin{align*} \begin{bmatrix} x_0^{(5)} \\ x_1^{(5)} \\ x_2^{(5)} \end{bmatrix} = \begin{bmatrix} 1 \\ 13 \\ 11 \end{bmatrix} \qquad \text{ requires } \qquad 13w_1 + 11w_2 \geq 150 \end{align*}\]

Combining all of the constraints would show the feasible region, i.e., the set of weights in the weight space that would correctly classify all training examples. If this region is non-empty, the training set is said to be linearly separable.

Definition: The feasible region is the intersection of all these constraints. It is the set of weights that correctly classify all training examples.

Definition: A dataset is linearly separable if there exists a hyperplane that perfectly separates the two classes. Equivalently, the feasible region is non-empty.

The visualization below is similar to the previous one, but has been updated to show the constraints that a data point imposes on the weight space while fixing \(w_0 = -150\). Pick a data point on the left to see the corresponding constraints on the right.

Weight-space constraints from data points

Figure 2: Interactive Pick a data point on the left to see the corresponding constraint in weight space on the right. Shown for linear classification with \(w_0 = -150\) fixed. Each data point defines a half-space of weights that classify that point correctly.

Having understood that each data point imposes a constraint on the weights, our goal is to find a set of weights to classify all the data points correctly. This is equivalent to finding one point in the feasible region. More generally, the problem of finding the entire feasible region given the constraints is a constraint satisfaction problem. When all the constraints are linear (as in our case), we can find the feasible region by solving a linear programming problem. Linear programming is a well-studied problem in mathematics and can be solved efficiently using algorithms such as the simplex method or interior point methods. These methods are beyond the scope of this course.

In practice, most datasets are not linearly separable. A linear model will misclassify at least one data point for such a dataset. Therefore, we cannot aim to learn a model that classifies all the data points correctly. Instead, we will ask a different question: how can we learn a model with “good” weights even when those weights do not classify every training point correctly? To make “good” precise, we need a loss function that measures how badly a set of weights fits the training data. We can then adjust the weights to reduce that loss.

3.3 Deriving a Decision Boundary

Although this binary linear classification model is not practical for most real-world datasets, it provides a useful framework for us to start exploring how to derive the decision boundary for a linear classification model. In this section, we will derive the expression for one decision boundary for a given linear classification model.

Let’s simplify the leaf prediction problem by considering three leaf data points in the lower left of the feature space. There are two oaks and one maple:

Table 2: Three leaf data points in the lower left of the feature space.
# Leaf Width Leaf Height Leaf Type
2 9.0 cm 6.0 cm Oak
4 5.0 cm 10.0 cm Oak
6 11.0 cm 9.0 cm Maple

Question: Each of these data points imposes a constraint on the weights. For the model to classify a data point correctly, the weights must satisfy an inequality. Derive the inequality constraint for each of the three leaves.

Answer:

The model predicts Oak (\(t = 0\)) when \(z < 0\) and Maple (\(t = 1\)) when \(z \geq 0\), where \(z = w_0 + w_1 x_1 + w_2 x_2\).

Oak leaf #2 has \(\begin{bmatrix} x_0^{(2)} & x_1^{(2)} & x_2^{(2)} \end{bmatrix}^\top = \begin{bmatrix} 1 & 9 & 6 \end{bmatrix}^\top\) and \(t^{(2)} = 0\). Correct classification requires \[\begin{align*} w_0 + 9 w_1 + 6 w_2 < 0. \end{align*}\]

Oak leaf #4 has \(\begin{bmatrix} x_0^{(4)} & x_1^{(4)} & x_2^{(4)} \end{bmatrix}^\top = \begin{bmatrix} 1 & 5 & 10 \end{bmatrix}^\top\) and \(t^{(4)} = 0\). Correct classification requires \[\begin{align*} w_0 + 5 w_1 + 10 w_2 < 0. \end{align*}\]

Maple leaf #6 has \(\begin{bmatrix} x_0^{(6)} & x_1^{(6)} & x_2^{(6)} \end{bmatrix}^\top = \begin{bmatrix} 1 & 11 & 9 \end{bmatrix}^\top\) and \(t^{(6)} = 1\). Correct classification requires \[\begin{align*} w_0 + 11 w_1 + 9 w_2 \geq 0. \end{align*}\]

It is feasible but cumbersome to derive a decision boundary by solving the system of inequalities. Instead, we will use a geometric approach to find the decision boundary.

Three leaf data points

Figure 3: Two oaks and one maple from the lower left of the original feature space.

Looking at the three points, a line representing the decision boundary should live somewhere in the gap between the two oak leaves and the maple leaf. Since the gap is fairly wide, there can be many such lines corresponding to multiple solutions to the system of inequalities. This means that we can choose the line that makes our calculations easier. Consider the line through the two points \(\begin{bmatrix} 14 & 0 \end{bmatrix}^\top\) and \(\begin{bmatrix} 0 & 14 \end{bmatrix}^\top\). This line passes through the gap and separates the two classes perfectly. It is also convenient for our calculations because it is at a 45-degree angle and corresponds to a slope of \(-1\).

Question: Derive the weight vector \(\begin{bmatrix} w_0 & w_1 & w_2 \end{bmatrix}^\top\) for the decision boundary given by the line through the two points \(\begin{bmatrix} 14 & 0 \end{bmatrix}^\top\) and \(\begin{bmatrix} 0 & 14 \end{bmatrix}^\top\).

Answer:

Reading from the plot, the line through the two points has slope \(-1\) and \(x_2\)-intercept \(14\). Therefore, its equation is \[\begin{align*} x_2 = -x_1 + 14. \end{align*}\] Moving all the quantities to the left side of the equation gives \[\begin{align*} x_1 + x_2 - 14 = 0. \end{align*}\] Comparing this equation with the decision boundary \(w_0 + w_1 x_1 + w_2 x_2 = 0\) suggests that the weights should be \[\begin{align*} \begin{bmatrix} w_0 & w_1 & w_2 \end{bmatrix}^\top = \begin{bmatrix} -14 & 1 & 1 \end{bmatrix}^\top. \end{align*}\]

However, it was an arbitrary decision for us to move all the quantities to the left side. Moving them to the right side instead gives the weights \[\begin{align*} \begin{bmatrix} w_0 & w_1 & w_2 \end{bmatrix}^\top = \begin{bmatrix} 14 & -1 & -1 \end{bmatrix}^\top, \end{align*}\] which describes the same line. To choose between the two, we plug in a labelled point and verify the sign of the prediction. The maple leaf #6 with \(\begin{bmatrix} x_0^{(6)} & x_1^{(6)} & x_2^{(6)} \end{bmatrix}^\top = \begin{bmatrix} 1 & 11 & 9 \end{bmatrix}^\top\) should produce a prediction of \(y = 1\). With the second set of weights, we get \[\begin{align*} z = \mathbf{w}^\top \mathbf{x}^{(6)} = \begin{bmatrix} 14 & -1 & -1 \end{bmatrix} \begin{bmatrix} 1 \\ 11 \\ 9 \end{bmatrix} = 14 - 11 - 9 = -6 < 0 \Rightarrow y = 0. \end{align*}\] This is the wrong prediction. With the first set of weights, we get \[\begin{align*} z = \mathbf{w}^\top \mathbf{x}^{(6)} = \begin{bmatrix} -14 & 1 & 1 \end{bmatrix} \begin{bmatrix} 1 \\ 11 \\ 9 \end{bmatrix} = -14 + 11 + 9 = 6 \geq 0 \Rightarrow y = 1. \end{align*}\] This is the correct prediction, so the correct weights are \[\begin{align*} \begin{bmatrix} w_0 & w_1 & w_2 \end{bmatrix}^\top = \begin{bmatrix} -14 & 1 & 1 \end{bmatrix}^\top. \end{align*}\]

The process of deriving the weights for a decision boundary highlights two useful strategies. First, geometric intuition can simplify mathematical derivations, an idea closely related to Fundamental Idea #4 (ML Describes Geometric Processes). Because we visually chose a decision boundary that simplified our calculations, we could easily write down the corresponding mathematical expression by inspecting the plot. Second, reasoning from first principles can complement geometric intuition. When determining the signs of the weights, we plugged in a labelled point and checked the sign of the prediction. This approach reminds us that whenever intuition is insufficient, we can always fall back on reasoning from first principles — in this case, testing our ideas by plugging in a labelled point.

4 Failed Approaches to Linear Classification

Before we present the logistic regression model, let’s assume that we are machine learning scientists and embark on a journey through several simpler approaches to designing such a model. Since these approaches are not the final logistic regression model, the purpose of this journey is to understand how each model works and analyze why it fails. These failed attempts motivate the careful design of logistic regression and help us appreciate the particular design choices that make it successful.

Recall that our journey starts with the linear classification model below. \[\begin{align*} z = \mathbf{w}^{\top} \mathbf{x} \qquad\qquad y = \begin{cases} 1 & \text{if } z \geq 0 \\ 0 & \text{if } z < 0 \end{cases} \end{align*}\] Since we cannot hope for the model to classify all the data points correctly, we would like to define a loss (or cost) function that measures how badly a set of weights fits the training data. The loss function measures the error for one example, and the cost is the average loss over the training set. Our goal is then to find a set of weights that minimizes the cost function. So the question is, which loss function should we use?

4.1 Why Not Use Zero-One Loss?

Since we care about classifying as many data points correctly as possible, the most natural choice for the loss/cost function is to count the mistakes. This idea gives us the zero-one loss function shown below. \[\begin{align*} \mathcal{L}_{0-1}(y, t) = \begin{cases} 0 & \text{if } y = t \\ 1 & \text{otherwise} \end{cases} = \mathbb{I}[y \neq t] \end{align*}\] where \(\mathbb{I}[e]\) is the indicator variable that evaluates to \(1\) if the condition \(e\) holds true, and \(0\) otherwise.

This loss function perfectly captures what we care about. The corresponding cost function counts the number of misclassified examples, which is the classification error.

The resulting model is shown below: \[\begin{align*} & \text{Linear model: } & & z = \mathbf{w}^{\top} \mathbf{x} \\[5pt] & \text{Determine prediction: } & & y = \begin{cases} 1 & \text{if } z \geq 0 \\ 0 & \text{if } z < 0 \end{cases} \\[5pt] & \text{Calculate loss: } & & \mathcal{L}_{0-1}(y, t) = \begin{cases} 0 & \text{if } y = t \\ 1 & \text{otherwise} \end{cases} = \mathbb{I}[y \neq t] \end{align*}\]

We already told you that this model would fail. Let’s see exactly how it fails. To train this model with gradient descent, we need the gradient of the loss with respect to the weights. The plot below shows how the loss of a negative example varies with \(z\). It is a step function that is \(0\) when the prediction is correct and \(1\) when it is incorrect.

Zero-one loss for a negative example

Figure 4: As a function of \(z\) for a negative example (\(t = 0\)), the loss is 0 when \(z < 0\) (correct prediction) and 1 when \(z \geq 0\) (incorrect prediction).

Question: Consider a negative example with the feature vector \(\mathbf{x} = \begin{bmatrix} 1 & -2 \end{bmatrix}^\top\) and the target \(t = 0\), where the first entry of \(\mathbf{x}\) is the dummy feature for the bias. Compute the gradient of the zero-one loss with respect to the weights for each of the three weight vectors below.

Recall that we can compute the gradient using the chain rule shown below. You may also want to consult Figure 4 above. \[\begin{align*} \frac{\partial \mathcal{L}_{0-1}}{\partial w_j} = \frac{\partial \mathcal{L}_{0-1}}{\partial z} \frac{\partial z}{\partial w_j} \end{align*}\]

\(\mathbf{w}\) \(z\) \(\frac{\partial \mathcal{L}_{0-1}}{\partial w_j}\)
\(\begin{bmatrix} 1 & 1 \end{bmatrix}^\top\)
\(\begin{bmatrix} 2 & 1 \end{bmatrix}^\top\)
\(\begin{bmatrix} 3 & 1 \end{bmatrix}^\top\)

Answer:

Based on the chain rule, the gradient of the loss with respect to each weight depends on the partial derivative of the loss with respect to \(z\). Since we are given the feature vector \(\mathbf{x} = \begin{bmatrix} 1 & -2 \end{bmatrix}^\top\), we can compute the \(z\) value for each weight vector as follows. \[\begin{align*} z = \begin{bmatrix} w_0 & w_1 \end{bmatrix} \begin{bmatrix} 1 \\ -2 \end{bmatrix} = w_0 - 2 w_1. \end{align*}\]

From Figure 4, we notice that the loss is a flat line for all values of \(z\) except at \(z = 0\), where it jumps. This means that the partial derivative of the loss with respect to \(z\) is \(0\) everywhere except at \(z = 0\), where it is undefined. By the chain rule, the gradient of the loss behaves the same way. We get the completed table below.

\(\mathbf{w}\) \(z\) \(\frac{\partial \mathcal{L}_{0-1}}{\partial w_j}\)
\(\begin{bmatrix} 1 & 1 \end{bmatrix}^\top\) \(1 - 2 = -1\) \(0\)
\(\begin{bmatrix} 2 & 1 \end{bmatrix}^\top\) \(2 - 2 = 0\) undefined
\(\begin{bmatrix} 3 & 1 \end{bmatrix}^\top\) \(3 - 2 = 1\) \(0\)
The first and third weight vectors both result in a gradient of \(0\), even though the third weight vector misclassifies the example. The gradient does not give us useful signals about how to improve the weights.

Let’s summarize our observations. We want to use gradient descent to find the weights that minimize the cost for a training set. This means we need a loss function that produces useful gradient signals for model training. This implies two minimum requirements. One, the loss function should be differentiable everywhere. The zero-one loss fails this requirement at \(z = 0\). Two, the loss function should produce non-zero gradients for all points. The zero-one loss fails this requirement for almost all points. In the next section, we will explore a loss function that we have encountered before that meets these two requirements.

4.2 Why Not Use Squared Error Loss?

Recall that we successfully used the squared error loss function for linear regression in the previous chapter. Moreover, the squared error loss meets the two minimum requirements proposed above. It is differentiable everywhere and produces non-zero gradients for many points. For this attempt, let’s experiment with using the squared error loss for linear classification.

Adding the squared error loss to the linear model gives the model shown below: \[\begin{align*} & \text{Linear model: } & & z = \mathbf{w}^{\top} \mathbf{x} \\[5pt] & \text{Determine prediction: } & & y = \begin{cases} 1 & \text{if } z \geq 0.5 \\ 0 & \text{if } z < 0.5 \end{cases} \\[5pt] & \text{Calculate loss: } & & \mathcal{L}_{SE}(z, t) = \frac{1}{2}(z - t)^2 \end{align*}\]

Two details in this model are worth discussing. First, the threshold is changed from \(0\) to \(0.5\). Because \(z\) directly affects the loss, a negative example drives \(z\) towards \(0\) and a positive example drives \(z\) towards \(1\). Therefore, the natural place for the threshold is \(0.5\), the midpoint of the two targets. Second, the loss function uses the intermediate value \(z\) rather than the prediction \(y\). If we change the squared error loss to use \(y\) instead, the resulting model is equivalent to the previous model using the zero-one loss. We leave this as an exercise for you.

Question: Suppose we change the squared error loss to use the prediction \(y\) instead of the intermediate value \(z\), giving \(\mathcal{L}_{SE}(y, t) = \frac{1}{2}(y - t)^2\). Show that the resulting model is equivalent to the model with the zero-one loss from the previous section.

While this approach seems reasonable, it has a fundamental problem. The squared error loss can be large even when the prediction is correct! Let’s analyze this problem with some concrete examples.

Question: Consider a positive example with the target \(t = 1\). For each value of \(z\) in the table below, determine the prediction \(y\), whether the prediction is correct, and the squared error loss.

\(t\) \(z\) Prediction \(y\) Correct or wrong? \(\mathcal{L}_{SE}\)
\(1\) \(1\)
\(1\) \(10\)
\(1\) \(100\)

Summarize your observations by circling the best choice in the sentence below.

For a positive example with \(t = 1\), when \(z\) is large (e.g., \(z = 10\) or \(z = 100\)), the prediction is (correct / wrong) but the squared error loss is (small / large).

Answer:

\(t\) \(z\) Prediction \(y\) Correct or wrong? \(\mathcal{L}_{SE}\)
\(1\) \(1\) \(1\) Correct \(\frac{1}{2}(1 - 1)^2 = 0\)
\(1\) \(10\) \(1\) Correct \(\frac{1}{2}(10 - 1)^2 = 40.5\)
\(1\) \(100\) \(1\) Correct \(\frac{1}{2}(100 - 1)^2 = 4900.5\)

Summary of observations: For a positive example with \(t = 1\), when \(z\) is large (e.g., \(z = 10\) or \(z = 100\)), the prediction is correct but the squared error loss is large.

The model classifies all three points correctly. However, the loss grows rapidly as \(z\) moves further above the target value of \(1\). During gradient descent, this loss would push the weights to decrease \(z\) toward \(1\), penalizing a prediction that is correct and increasingly confident.

We can also understand this problem by looking at the visualization below. The blue line shows how \(z\) changes monotonically with the feature \(x\). Although a large positive \(z\) value produces the correct prediction for a positive example, the squared error loss would penalize the difference \(z - t\) heavily. You could interpret \(z\) as a measure of how confident the model is in its prediction. With this interpretation, it is even more confusing that the model would penalize a prediction when the prediction is correct and made with high confidence.

Squared error for classification

Figure 5: The squared error loss penalizes a correct prediction. The dots are the targets, with blue at \(t = 0\) and orange at \(t = 1\). The blue line is the model output \(z\), and the dashed line is the threshold at \(0.5\). Every dot is on the correct side of the threshold, so all six examples are classified correctly. But the loss measures how far \(z\) is from the target, shown by the orange bar at \(x = 15\). That example has \(z \approx 3\) against a target of \(t = 1\), so it incurs a large loss despite being correct.

The root of the problem is a mismatch between what the squared error loss measures and what we aim to optimize in classification. By treating \(z\) as an estimate of the target, the squared error measures the difference between \(z\) and \(t\) and is minimized when they are equal. Classification only cares about whether \(z\) is on the correct side of the threshold. Once \(z\) is past the threshold, a large \(z\) can be interpreted as the model being more confident, and confidence should not be penalized.

What makes this mismatch worse is that \(z\) is unbounded. If \(z\) grows without limit, so can the difference \(z - t\) and the loss. This is not a problem in regression, where we care about making \(z\) match the target \(t\).

This analysis suggests a potential solution. Since our targets only take the values \(0\) and \(1\), the value we compare the target against should live in the same range. What if we apply a function to \(z\) that restricts it to the interval \([0, 1]\)? This way, the difference between \(z\) and \(t\) can never exceed \(1\), so confidence cannot produce a large loss. In the next section, we will introduce the sigmoid function, which we use to restrict the prediction to the interval \([0, 1]\).

4.3 Introducing the Sigmoid Function

Our next attempt will not involve replacing our loss function. Instead, we will add the sigmoid function as an intermediate step of the model to restrict the prediction to the interval \([0,1]\). To prepare for this attempt, we will focus on defining the sigmoid function and discussing its important properties in this section.

Definition: The sigmoid function (also called the logistic function) is defined as: \[\begin{align*} \sigma(z) = \frac{1}{1 + e^{-z}} \end{align*}\]

The name “sigmoid” refers to the function’s S-shaped curve. Strictly speaking, it names a whole family of S-shaped functions, and the function above is the logistic function. However, we will use “sigmoid” since it is the convention in machine learning.

Definition: The value \(z = \mathbf{w}^{\top} \mathbf{x}\) is called the pre-activation or logit.

The term “pre-activation” comes from the neural network literature. In a neural network, each unit computes the pre-activation \(z\) before passing it through an activation function to produce the output.

The term “logit” comes from statistics. If \(y = \sigma(z)\) is interpreted as a probability of an event, then \(z\) represents the log-odds of the event: \(z = \log\left(\frac{y}{1-y}\right)\).

The sigmoid function has several important properties that make it suitable for binary classification. The output is always in the interval \((0, 1)\), which allows us to interpret it as a probability. The function is strictly increasing, meaning that larger values of \(z\) correspond to higher probabilities of the positive class. The sigmoid satisfies a symmetry property: \(\sigma(-z) = 1 - \sigma(z)\), reflecting the symmetry between the two classes.

The sigmoid function behaves intuitively at specific values. When \(z = 0\), we have \(\sigma(0) = 0.5\), indicating maximum uncertainty between the two classes. As \(z\) becomes large and positive (\(z \to \infty\)), \(\sigma(z)\) approaches 1, indicating high confidence in the positive class. Conversely, as \(z\) becomes large and negative (\(z \to -\infty\)), \(\sigma(z)\) approaches 0, indicating high confidence in the negative class. Generally, we will treat \(\sigma(z) = 0.5\) as the decision boundary.

Sigmoid activation function

Figure 6: \(\sigma(z) = \frac{1}{1 + e^{-z}}\).

The derivative of \(\sigma(z)\) with respect to \(z\) has an elegant form. We leave the derivation as an exercise below:

\[\begin{align*} \frac{d\sigma(z)}{dz} = \sigma(z)(1 - \sigma(z)) \end{align*}\]

Question: Derive the derivative of the sigmoid function from its definition.

\[\sigma(z) = \frac{1}{1 + e^{-z}} \qquad \Rightarrow \qquad \frac{d\sigma(z)}{dz} = \sigma(z)(1 - \sigma(z)).\]

Answer:

Write the sigmoid as \(\sigma(z) = (1 + e^{-z})^{-1}\) and apply the chain rule. \[\begin{align*} \frac{d\sigma(z)}{dz} &= -(1 + e^{-z})^{-2} \cdot \frac{d}{dz}\left(1 + e^{-z}\right) \\[5pt] &= -(1 + e^{-z})^{-2} \cdot \left(-e^{-z}\right) \\[5pt] &= \frac{e^{-z}}{(1 + e^{-z})^2} \end{align*}\]

This expression is correct but does not yet resemble the version we want. Split the expression into two factors, the first of which is \(\sigma(z)\). \[\begin{align*} \frac{e^{-z}}{(1 + e^{-z})^2} = \frac{1}{1 + e^{-z}} \cdot \frac{e^{-z}}{1 + e^{-z}} = \sigma(z) \cdot \frac{e^{-z}}{1 + e^{-z}} \end{align*}\]

Next, rewrite the second factor into the desired form. \[\begin{align*} \frac{e^{-z}}{1 + e^{-z}} = \frac{(1 + e^{-z}) - 1}{1 + e^{-z}} = 1 - \frac{1}{1 + e^{-z}} = 1 - \sigma(z) \end{align*}\]

Combining the two factors gives the desired version. \[\begin{align*} \frac{d\sigma(z)}{dz} = \sigma(z)(1 - \sigma(z)) \end{align*}\]

This derivative will greatly simplify the gradient computation for training logistic regression, which we will see shortly.

4.4 Why Not Sigmoid with Squared Error?

We are now ready to add the sigmoid function to our model with squared error loss. The resulting model is given below.

\[\begin{align*} & \text{Linear model: } & & z = \mathbf{w}^{\top} \mathbf{x} \\[5pt] & \text{Calculate probability: } & & y = \sigma(z) = \frac{1}{1 + e^{-z}} \\[5pt] & \text{Calculate loss: } & & \mathcal{L}_{SE}(y, t) = \frac{1}{2}(y - t)^2 \end{align*}\]

Unlike previous attempts, this combination has the advantage that both the activation function and the loss function are differentiable. We can compute the gradient using the chain rule: \[\begin{align*} \frac{\partial \mathcal{L}_{SE}}{\partial w_j} &= \frac{\partial \mathcal{L}_{SE}}{\partial y} \cdot \frac{dy}{dz} \cdot \frac{\partial z}{\partial w_j} = (y - t) \cdot y(1 - y) \cdot x_j \end{align*}\]

Unfortunately, this model suffers from a problem with gradient signal. We could run into scenarios where the prediction is very wrong but the weights appear to be at a critical point. Let’s consider some examples.

Question: Consider the model above with the sigmoid activation and the squared error loss. For each row in the table below, determine the approximate value of the prediction \(y\), whether the prediction is correct or wrong, and the gradient \(\frac{\partial \mathcal{L}_{SE}}{\partial w_j}\). Then decide whether the weights appear to be at a critical point.

You may want to consult Figure 6. Recall that the weights are at a critical point when the gradient is close to \(0\).

Target \(t\) \(z\) Prediction \(y\) Correct or wrong? Gradient \(\frac{\partial \mathcal{L}_{SE}}{\partial w_j}\) Critical point?
\(1\) \(-1000\)
\(0\) \(1000\)

Summarize your observations by circling the best choice in the next sentence. When the prediction is (correct / wrong / very wrong), the weights appear to be (at / not at) a critical point.

Answer:

For the first row, \(z\) is a large negative number. Reading off Figure 6, we can see that the prediction \(y\) is close to \(0\) since \(\sigma(-1000) \approx 0\). The target is \(t = 1\), so the prediction is very wrong.

What about the gradient? Recall the following expression for the gradient.

\[\begin{align*} \frac{\partial \mathcal{L}_{SE}}{\partial w_j} = \underbrace{(y - t)}_{\approx -1} \cdot \underbrace{y(1 - y)}_{\approx 0} \cdot x_j \end{align*}\]

The first factor \((y - t)\) is as large as it can be, correctly reporting that the prediction is far from the target. However, the second factor \(y(1 - y)\) is close to \(0\) at \(z = -1000\) and drives the entire gradient to \(0\). Therefore, the prediction is very wrong but the weights appear to be at a critical point.

We can complete the second row using similar reasoning. There, \(y = \sigma(1000) \approx 1\), so the first factor is \((y - t) \approx 1\) but the second factor \(y(1 - y)\) is close to \(0\) once again. See the complete table below.

Target \(t\) \(z\) Prediction \(y\) Correct or wrong? Gradient \(\frac{\partial \mathcal{L}_{SE}}{\partial w_j}\) Critical point?
\(1\) \(-1000\) \(\approx 0\) Very wrong \(\approx 0\) Yes
\(0\) \(1000\) \(\approx 1\) Very wrong \(\approx 0\) Yes

Summary of observations: When the prediction is very wrong, the weights appear to be at a critical point.

Gradient descent relies on the gradient to provide useful signals for learning. In both rows, the model badly needs to adjust its weights because the prediction is very wrong, but the weight update is tiny because the gradient is close to \(0\). The model is stuck precisely where it most needs to learn.

This attempt illustrates an important lesson: when building a machine learning architecture with modular components (e.g., the hypothesis, loss, and optimizer), the choice of each component is not independent. Combining the sigmoid activation with squared error results in an unfortunate interaction: the very shape of the sigmoid that bounds outputs to \((0, 1)\) is also what shrinks the gradient to near zero at the extremes. Understanding the underlying mathematics is therefore essential, and we cannot simply mix and match components without checking how they interact. We need a loss function whose gradient remains strong even when the model is confidently wrong.

5 Logistic Regression

Having seen why these previous approaches fail, we now develop the complete logistic regression model.

5.1 The Cross-Entropy Loss

The key insight is to combine the sigmoid activation function with a specially designed loss function, the cross-entropy loss, that provides good gradient signal throughout the training process.

Definition: The cross-entropy loss for binary classification is defined as: \[\begin{align*} \mathcal{L}_{CE}(y, t) & = \begin{cases} -\log(y), \quad\qquad \text{ if } t = 1 \\[3pt] -\log(1-y), \quad \text{ if } t = 0 \end{cases} \\[5pt] & = -t \log(y) - (1-t) \log(1-y) \\ \end{align*}\]

where \(y\) is the predicted probability of the positive class and \(t \in \{0, 1\}\) is the target.

This loss function can be understood in two equivalent ways. First, when \(t = 1\), the loss reduces to \(-\log(y)\), which is large when \(y\) is close to 0 (confident wrong prediction) and small when \(y\) is close to 1 (confident correct prediction). Similarly, when \(t = 0\), the loss reduces to \(-\log(1-y)\), which is large when \(y\) is close to 1 and small when \(y\) is close to 0.

Second, the cross-entropy loss has a probabilistic interpretation. If we interpret \(y\) as the probability that \(t = 1\), then the loss measures the negative log-likelihood of the observed target under the model’s predicted distribution. Minimizing this loss corresponds to maximum likelihood estimation, which is a statistical method we will explore in Chapter 7.

Cross-entropy versus squared error

Figure 7: For a positive example (\(t = 1\)), the cross-entropy loss (blue) increases steeply as the prediction moves away from the target, while squared error (orange) grows quadratically. The cross-entropy provides stronger gradient signal for very wrong predictions.

The key advantage of cross-entropy loss becomes clear when we compute its gradient. Combining the sigmoid activation with the cross-entropy loss gives the model shown below.

\[\begin{align*} & \text{Linear model: } & & z = \mathbf{w}^{\top} \mathbf{x} \\[5pt] & \text{Calculate prediction: } & & y = \sigma(z) = \frac{1}{1 + e^{-z}} \\[5pt] & \text{Calculate loss: } & & \mathcal{L}_{CE}(y, t) = -t \log(y) - (1-t) \log(1-y) \end{align*}\]

Question: For the model above, compute the gradient of the cross-entropy loss with respect to the weight \(w_j\). Show that it simplifies to \[\begin{align*} \frac{\partial \mathcal{L}_{CE}}{\partial w_j} &= (y - t) x_j \end{align*}\]

You may want to use the chain rule and the derivative of the sigmoid function below. \[\frac{d\sigma(z)}{dz} = \sigma(z)(1 - \sigma(z)).\]

Answer:

The weight \(w_j\) affects the loss through \(z\) and then through \(y\), so we apply the chain rule: \[\begin{align*} \frac{\partial \mathcal{L}_{CE}}{\partial w_j} &= \frac{\partial \mathcal{L}_{CE}}{\partial y} \cdot \frac{dy}{dz} \cdot \frac{\partial z}{\partial w_j} \end{align*}\]

Each term can be computed in turn:

\[\begin{align*} \frac{\partial \mathcal{L}_{CE}}{\partial y} &= -\frac{t}{y} + \frac{1-t}{1-y} = \frac{-t+y}{y(1-y)}\\ \frac{dy}{dz} &= y(1-y) \\ \frac{\partial z}{\partial w_j} &= x_j \end{align*}\]

Combining these terms, we have \[\begin{align*} \frac{\partial \mathcal{L}_{CE}}{\partial w_j} &= \frac{-t+y}{y(1-y)} \cdot y(1-y) \cdot x_j \\ &= (-t + y) \cdot x_j \\ &= (y - t) x_j \end{align*}\]

The gradient vector has the following form: \[\begin{align*} \nabla_{\mathbf{w}} \mathcal{L}_{CE} &= (y - t) \mathbf{x} \end{align*}\]

Notice that the denominator of \(\frac{\partial \mathcal{L}_{CE}}{\partial y}\) cancels out with the \(y(1-y)\) term from the sigmoid derivative. This cancellation gives us two advantages. One, the gradient has an extremely simple final form. Two, removing the \(y(1-y)\) factor prevents the gradient from vanishing when \(y\) is close to \(0\) or \(1\), the exact problem we struggled with in the previous attempt. What remains is the error \((y - t)\), so the size of the gradient now reflects how wrong the prediction is. When the prediction \(y\) is far from the target \(t\), the gradient \((y - t) x_j\) stays large in magnitude, encouraging the model to make large weight updates. This is precisely the behaviour we wished to achieve.

5.2 The Complete Model

We can now define the complete logistic regression model by combining all the pieces.

Definition: The logistic regression model for binary classification has the hypothesis form \[\begin{align*} & z = \mathbf{w}^{\top} \mathbf{x} \\[5pt] & y = \sigma(z) = \frac{1}{1 + e^{-z}} \\[5pt] & \mathcal{L}(y, t) = -t \log(y) - (1-t) \log(1-y) \end{align*}\]

where

  • \(\mathbf{x} \in \mathbb{R}^{D+1}\) is the input vector for \(D\) features, which includes a dummy feature \(x_0 = 1\),
  • \(\mathbf{w} \in \mathbb{R}^{D+1}\) is the weight vector, which includes the bias term,
  • \(z\) is an intermediate real-valued quantity that we compute,
  • \(\sigma\) is the sigmoid function,
  • \(y \in (0, 1)\) is the predicted probability of the positive class,
  • \(t \in \{0, 1\}\) is the target label, and
  • \(\mathcal{L}(y, t)\) is the cross-entropy loss.

5.3 Cost Function and Training

To train the logistic regression model, we find the weights that minimize the average loss over the entire training set. That is, our cost function (also called the objective function) can be written as follows: \[\begin{align*} \mathcal{E}(\mathbf{w}) &= \frac{1}{N} \sum_{i=1}^N \mathcal{L}(y^{(i)}, t^{(i)}) \\ &= -\frac{1}{N} \sum_{i=1}^N \left[ t^{(i)} \log(y^{(i)}) + (1-t^{(i)}) \log(1-y^{(i)}) \right] \end{align*}\] where \(N\) is the number of training examples, and \(y^{(i)} = \sigma(\mathbf{w}^{\top} \mathbf{x}^{(i)})\) is the model’s prediction for the \(i\)-th training example.

To minimize this cost function, we can continue to use gradient descent. We can again use the linearity of gradients and apply the result from the previous section: \[\begin{align*} \nabla_{\mathbf{w}} \mathcal{E}(\mathbf{w}) &= \frac{1}{N} \sum_{i=1}^N \nabla_{\mathbf{w}} \mathcal{L}(y^{(i)}, t^{(i)}) \\ &= \frac{1}{N} \sum_{i=1}^N (y^{(i)} - t^{(i)}) \mathbf{x}^{(i)} \\ \end{align*}\]

The gradient descent update rule for the weights is: \[\begin{align*} \mathbf{w} \leftarrow \mathbf{w} - \alpha \frac{\partial \mathcal{E}}{\partial \mathbf{w}} = \mathbf{w} - \frac{\alpha}{N} \sum_{i=1}^N (y^{(i)} - t^{(i)}) \mathbf{x}^{(i)} \end{align*}\]

where \(\alpha\) is the learning rate.

Does this update rule look familiar to you? It has exactly the same form as the gradient descent update for linear regression with the squared error loss! The only difference is in how \(y^{(i)}\) is computed: in linear regression, \(y^{(i)} = \mathbf{w}^{\top} \mathbf{x}^{(i)}\), while in logistic regression, \(y^{(i)} = \sigma(\mathbf{w}^{\top} \mathbf{x}^{(i)})\). This similarity is not a coincidence. Both models belong to a family known as generalized linear models, in which the activation function and the loss function are paired in a particular way. For every model in this family, the gradient takes the form (prediction \(-\) target) \(\times\) input. Exploring the connections among these models is beyond the scope of this course.

Question: The update rule above sums over all \(N\) training examples. Rewrite it in vectorized form, so that matrix and vector operations replace the summation. That is, show that \[\begin{align*} \mathbf{w} \leftarrow \mathbf{w} - \frac{\alpha}{N} \sum_{i=1}^N (y^{(i)} - t^{(i)}) \mathbf{x}^{(i)} \Rightarrow \mathbf{w} \leftarrow \mathbf{w} - \frac{\alpha}{N} \mathbf{X}^\top (\mathbf{y} - \mathbf{t}) \end{align*}\]

You may use the following notation:

  • The data matrix \(\mathbf{X} \in \mathbb{R}^{N \times (D+1)}\) contains the feature vectors, where row \(i\) is \((\mathbf{x}^{(i)})^\top\),
  • the target vector \(\mathbf{t} \in \mathbb{R}^{N \times 1}\) contains the targets, and
  • the prediction vector \(\mathbf{y} \in \mathbb{R}^{N \times 1}\) contains the predictions.

Answer:

Consider the summation. Each term is a scalar \((y^{(i)} - t^{(i)})\) multiplied by a vector \(\mathbf{x}^{(i)} \in \mathbb{R}^{(D+1) \times 1}\). Stacking the scalars gives the residual vector \(\mathbf{r} = \mathbf{y} - \mathbf{t} \in \mathbb{R}^{N \times 1}\), and stacking the vectors gives the data matrix \(\mathbf{X} \in \mathbb{R}^{N \times (D+1)}\).

The gradient vector must be a column vector in \(\mathbb{R}^{(D+1) \times 1}\), with one derivative for each of the \((D+1)\) weights. The dimension over \(N\) should disappear after the product, so the only valid choice is \(\mathbf{X}^\top \mathbf{r}\). \[\begin{align*} \nabla_{\mathbf{w}} \mathcal{E}(\mathbf{w}) = \frac{1}{N} \sum_{i=1}^N (y^{(i)} - t^{(i)}) \mathbf{x}^{(i)} = \frac{1}{N} \mathbf{X}^\top (\mathbf{y} - \mathbf{t}) \end{align*}\]

Substituting this expression into the update rule, we have \[\begin{align*} \mathbf{w} \leftarrow \mathbf{w} - \frac{\alpha}{N} \mathbf{X}^\top \left( \mathbf{y} - \mathbf{t} \right) \end{align*}\]

5.4 Decision Boundaries

Understanding the decision boundaries of logistic regression provides geometric intuition about how the model makes predictions. Recall that in logistic regression, the predicted probability is \(y = \sigma(\mathbf{w}^{\top} \mathbf{x})\). To make a binary prediction, we typically use a threshold of 0.5: predict 1 if \(y \geq 0.5\), and 0 otherwise.

Since \(\sigma(0) = 0.5\), this is equivalent to predicting 1 when \(\mathbf{w}^{\top} \mathbf{x} \geq 0\) and 0 when \(\mathbf{w}^{\top} \mathbf{x} < 0\). Thus, the decision boundary is the set of points where \(\mathbf{w}^{\top} \mathbf{x} = 0\). This is a hyperplane in the input space, just as in the linear classification model we started with.

5.5 Feature Scaling

Just like in linear regression, the gradient descent optimization process used for logistic regression can be sensitive to the scale of features. It is generally recommended to standardize features to have zero mean and unit variance before training: \[\begin{align*} \tilde{x}_j = \frac{x_j - \mu_j}{\sigma_j} \end{align*}\] where \(\mu_j\) and \(\sigma_j\) are the mean and standard deviation of feature \(j\) computed from the training set.

6 Comparison with k-Nearest Neighbours

It is instructive to compare logistic regression with the k-nearest neighbours (k-NN) algorithm studied in the first chapter. These two algorithms represent different approaches to classification, each with distinct advantages and limitations.

Logistic regression is a parametric model. The hypothesis space consists of all functions that can be represented using a fixed number of parameters (the weights \(\mathbf{w}\)). Once we have learned the weights from the training data, we can discard the training data itself. The model is completely described by the parameter vector \(\mathbf{w} \in \mathbb{R}^{D+1}\), where \(D\) is the number of features and the extra parameter is the bias.

In contrast, k-NN is a non-parametric model. The hypothesis space cannot be described using a fixed number of parameters. Instead, the model is defined directly in terms of the training data. To make a prediction, we need to store all training examples and compute distances to them. In a sense, non-parametric models have an effectively infinite number of parameters that grow with the amount of training data.

Definition: A machine learning model is parametric if its hypothesis space can be characterized by a fixed-size set of parameters. A model is non-parametric if the complexity of the hypothesis can grow with the amount of training data.

The choice between parametric and non-parametric models involves tradeoffs along several dimensions.

The first dimension is the inductive bias of each model, a concept that is central to understanding these tradeoffs.

Definition: The inductive bias of a machine learning algorithm is the set of assumptions or preferences that guide how it generalizes from training data to new, unseen examples. These assumptions determine what patterns the algorithm can learn and how it interpolates between training points.

The inductive bias of logistic regression is that the classes can be separated by a linear decision boundary, that is, a hyperplane. The inductive bias of k-NN is that similar inputs should have similar outputs, so it can represent arbitrarily complex boundaries given enough data. A strong assumption helps when it matches the data and hurts when it does not.

The two models also differ in storage and prediction cost. Logistic regression stores \(D+1\) weights and predicts by computing \(\mathbf{w}^{\top} \mathbf{x}\) and applying the sigmoid function, no matter how large the training set is. k-NN stores every training example and computes distances to all of them for each prediction, which is slow for large datasets. This difference follows directly from the parametric distinction above: logistic regression discards the training data after fitting, while k-NN cannot.

Next, consider how interpretable each model is. The weights of logistic regression tell us which features matter most and in which direction. With k-NN, an individual prediction is easy to explain, since we can point to the neighbours responsible for it, but the behaviour of the model as a whole is harder to characterize.

Finally, the two models generalize differently. Logistic regression generalizes well from limited data when the linear assumption holds, and underfits when it does not. k-NN adapts to complex patterns, but overfits to noise when \(k\) is too small and degrades in high-dimensional spaces due to the curse of dimensionality.

The table below summarizes these tradeoffs.

Table 3: Tradeoffs between logistic regression and k-NN.
Logistic Regression k-Nearest Neighbours
Model type Parametric Non-parametric
Inductive bias Linear decision boundary Nearby inputs share labels
Parameters \(D+1\) weights Grows with the data
Storage after training \(D+1\) weights Entire training set
Prediction cost One dot product Distance to every example
Interpretability Weights rank the features Neighbours explain one prediction
Typical failure Underfits a nonlinear boundary Overfits when \(k\) is small

Neither approach is universally superior. The choice depends on the specific problem, the amount and quality of available data, the computational constraints, and whether the inductive bias matches the underlying pattern in the data. In practice, logistic regression is often a good starting point due to its simplicity, efficiency, and interpretability. More complex models can be considered if the linear assumption is too restrictive.

7 Summary

Logistic regression is a fundamental algorithm for binary classification that elegantly combines a linear model structure with a probabilistic interpretation. The key design choices include using the sigmoid activation function and cross-entropy loss. While we motivated these choices through practical considerations about gradient signal and differentiability, they also reflect Fundamental Idea #5 (ML demands a probabilistic lens). The sigmoid output is a probability, and minimizing the cross-entropy loss is maximum likelihood estimation under a probabilistic model of binary outcomes. We will return to these statistical principles in Chapter 7.

The derivation in this chapter is an extended illustration of Fundamental Idea #1 (learning is optimization). Treating learning as optimization allows us to mix and match a model, a loss function, and an optimizer. However, our failed attempts showed that these components interact: a loss function is only useful if it provides the optimizer with a usable gradient.

As a parametric model, logistic regression makes strong assumptions about the form of the decision boundary. Here we see Fundamental Idea #4 (ML describes geometric processes) at work: the model splits the data space into two half-spaces, and training searches the weight space for one that separates the two classes. This inductive bias can help the model generalize from limited data, but it also limits the model’s expressiveness. Understanding this tradeoff helps us choose appropriate models for different problems and motivates extensions that we will discuss in future chapters.

Logistic regression remains one of the most widely used machine learning algorithms, valued for its simplicity, efficiency, and interpretability. It serves as an excellent baseline for binary classification tasks and as a building block for more complex models. The principles underlying logistic regression form the foundation for deep learning and neural networks, which we will explore in future chapters.