Multi-Class Classification

Learning Objectives

After reading this page, you should be able to:

  1. Explain why one-hot vectors are better than integers for encoding multi-class targets.
  2. Analyze desirable properties of the softmax function.
  3. Define the softmax regression model.
  4. Determine the decision boundaries of a softmax regression model with three classes.

1 Introduction

The previous chapter focused on a binary classification task, in which we only had to pick between two possible classes. In this chapter, we will extend this concept to multi-class classification, where each input example belongs to exactly one of \(K\) possible classes and \(K \geq 3\). This setting generalizes binary classification, which only distinguishes between two labels (\(K = 2\)).

Multi-class classification problems arise naturally in practice, such as:

  • Recognizing handwritten digits, where the classes correspond to the digits 0 through 9 (\(K = 10\)).
  • Identifying the species of an animal from an image.

1.1 Leaf Classification with Three Classes

Let’s extend our leaf classification example from previous chapters to a multi-class classification task. Each leaf is still described by two measurements: width and height. But now our task is to predict whether a leaf is an Oak (🍂), Maple (🍁), or Birch (🌿). Our training data consists of ten labelled examples:

Table 1: Training data for the multi-class leaf prediction problem.
# Leaf Width Leaf Height Leaf Type
1 7.0 cm 12.0 cm 🍂
2 9.0 cm 6.0 cm 🍂
3 10.0 cm 18.0 cm 🍂
4 5.0 cm 10.0 cm 🍂
5 13.0 cm 11.0 cm 🍁
6 11.0 cm 9.0 cm 🍁
7 9.0 cm 14.0 cm 🍁
8 6.0 cm 9.0 cm 🌿
9 7.5 cm 5.5 cm 🌿
10 10.0 cm 7.5 cm 🌿

Here is our data plotted in 2D space. Width is on the x-axis and height is on the y-axis.

Multi-class leaf training data

Figure 1: The example dataset is plotted by leaf width and height.

2 Formalizing the Problem

Before building a model, we need to write down our data precisely. In this section, we will develop a formal setup for the data, the targets, and the predictions with \(K \geq 3\) classes.

In logistic regression, we were given a dataset of \(N\) labelled examples: \[ \begin{aligned} \left\{ (\mathbf{x}^{(1)}, t^{(1)}), (\mathbf{x}^{(2)}, t^{(2)}), \dots, (\mathbf{x}^{(N)}, t^{(N)}) \right\} \end{aligned} \]

where \(\mathbf{x}^{(i)} \in \mathbb{R}^D\) is a feature vector and each target is a scalar, \(t^{(i)} \in \{0, 1\}\). With only two classes, one number is sufficient to distinguish the two classes. How should we encode the targets when we have at least three classes? A natural first attempt is to use an integer to denote each class. For example, Oak is \(1\), Maple is \(2\), and Birch is \(3\). This approach is quite general and can work for any number of classes.

Unfortunately, this approach is problematic because integer encodings imply an order and a distance that do not make sense for the class categories. For example, the encoding above may suggest that Maple lies “between” Oak and Birch. It is equally odd to assume that Oak is closer to Maple (\(|1 - 2| = 1\)) than to Birch (\(|1 - 3| = 2\)). A model trained with these target integers will treat these relationships as meaningful when they are not. What is worse is that the model’s behaviour would change if we simply renumbered the classes, which should make no difference.

A better approach is to encode each target as a one-hot vector, which is a vector in \(\mathbb{R}^K\) with a \(1\) in the entry for the correct class and \(0\) everywhere else. That is, exactly one entry is “hot” (\(=1\)). The target for an example in the \(k\)-th class can be denoted as follows, where the \(1\) is in the \(k\)-th row of the vector. \[ \begin{aligned} \mathbf{t}^{(i)} = \begin{bmatrix} 0 \\ \vdots \\ 1 \\ \vdots \\ 0 \end{bmatrix} \in \mathbb{R}^K \end{aligned} \]

For our leaf problem with \(K = 3\), the three classes can be encoded as: \[ \begin{aligned} \text{Oak: } \left[ \begin{matrix} 1 \\ 0 \\ 0 \end{matrix} \right] \qquad \text{Maple: } \left[ \begin{matrix} 0 \\ 1 \\ 0 \end{matrix} \right] \qquad \text{Birch: } \left[ \begin{matrix} 0 \\ 0 \\ 1 \end{matrix} \right] \end{aligned} \] For example, the fifth example in our dataset is a Maple leaf, so \(\mathbf{t}^{(5)} = [0, 1, 0]^\top\).

This one-hot encoding solves the two issues discussed above. First, it no longer assumes an order among the categories. No category has a higher numeric value than another. Second, the distance between any pair of one-hot vectors is the same.

Putting this together, in multi-class classification we are given a dataset: \[ \begin{aligned} \left\{ (\mathbf{x}^{(1)}, \mathbf{t}^{(1)}), (\mathbf{x}^{(2)}, \mathbf{t}^{(2)}), \dots, (\mathbf{x}^{(N)}, \mathbf{t}^{(N)}) \right\} \end{aligned} \] where \(\mathbf{x}^{(i)} \in \mathbb{R}^D\) is a feature vector as before, but each target \(\mathbf{t}^{(i)} \in \mathbb{R}^K\) is now a \(K\)-dimensional one-hot vector with an entry of \(1\) for the correct class.

3 Multi-Class Classification: Inference

The goal of inference is to answer the following question: given an input \(\mathbf{x}\) and a trained model, what class does the model predict? In this part, we describe how a multi-class classification model can map inputs to predictions. For now, we will focus on making predictions using a trained model and won’t worry about how the model is trained.

3.1 Prediction Vector

How can we represent our model’s predictions formally? Since there are \(K\) classes, we can represent the predictions in the same format as the targets. The predictions become a vector in \(\mathbb{R}^K\). \[ \begin{aligned} \mathbf{y} = \begin{bmatrix} y_1 \\ \vdots \\ y_k \\ \vdots \\ y_K \end{bmatrix} \in \mathbb{R}^K \end{aligned} \] Typically, \(\mathbf{y}\) forms a probability distribution over the \(K\) classes, satisfying the conditions below. \[ \begin{aligned} 0 \leq y_k \leq 1, \forall k = 1, \dots, K, \qquad \sum_{k=1}^K y_k = 1 \end{aligned} \]

3.2 Linear Scores (Logits)

The softmax regression model is a direct extension of the logistic regression model to at least three classes. To generate one prediction for each class, we calculate an intermediate score \(z_k \in \mathbb{R}\) for each class. This score is computed using a linear function of the features and the weights and biases.

Logistic regression required a weight vector to map the features to the prediction for one class. For softmax regression, we need to map \((D+1)\) features to \(K\) classes. As a result, softmax regression requires a weight matrix, not a vector.

Given an input vector \(\mathbf{x} \in \mathbb{R}^D\), the score for class \(k\) is computed as:

\[ \begin{aligned} z_k & = \sum_{j=1}^{D} w_{k,j} x_j + b_k \\ & = \left[ \begin{matrix} b_k & w_{k,1} & w_{k,2} & \dots & w_{k,D} \end{matrix} \right] \left[ \begin{matrix} 1 \\ x_1 \\ x_2 \\ \vdots \\ x_D \end{matrix} \right] \\ & = \mathbf{w}^{\top}_{k} \mathbf{x} \end{aligned} \] In the last line, we absorb the bias by adding a dummy feature \(1\) to the input, so \(\mathbf{w}_k\) and \(\mathbf{x}\) are both vectors in \(\mathbb{R}^{D+1}\).

We can collect the scores for all \(K\) classes in a single vector \(\mathbf{z} \in \mathbb{R}^K\) that can be calculated compactly as follows:

\[ \begin{aligned} \mathbf{z} & = \mathbf{W} \mathbf{x} \\ \mathbf{z} & = \left[ \begin{matrix} z_1 \\ z_2 \\ \vdots \\ z_K\end{matrix} \right] = \left[ \begin{matrix} b_1 & w_{1,1} & w_{1,2} & \dots & w_{1,D} \\ b_2 & w_{2,1} & w_{2,2} & \dots & w_{2,D} \\ \vdots & \vdots & \vdots & \ddots & \vdots \\ b_K & w_{K,1} & w_{K,2} & \dots & w_{K,D} \\ \end{matrix} \right] \left[ \begin{matrix} 1 \\ x_1 \\ x_2 \\ \vdots \\ x_D \end{matrix} \right] \end{aligned} \]

Note that our weight matrix \(\mathbf{W}\) has dimensions \(K \times (D + 1)\):

  • The \(k\)-th row in the matrix includes the parameters needed for calculating the \(k\)-th score \(z_k\).
  • Each row includes \(D + 1\) parameters (one weight for each feature and a bias term).

Classifying using these scores is straightforward. We simply pick the class corresponding to the largest score.

At this stage, the scores in \(\mathbf{z}\) are often referred to as logits; they are unnormalized and cannot yet be interpreted as probabilities.

3.3 Softmax Converts Scores to Probabilities

To convert logits into probabilities, we apply the softmax function. Just as the sigmoid maps a single logit to a probability, the softmax function maps a vector of logits to a probability distribution over the \(K\) classes.

Definition: Given a vector of logits \(\mathbf{z} \in \mathbb{R}^K\), where \(z_k\) is the logit for class \(k\), the softmax function produces a vector of predictions \(\mathbf{y} = \text{softmax}(\mathbf{z}) \in \mathbb{R}^K\), where the predicted probability for class \(k\) is \[ y_k = \frac{\exp(z_k)}{\sum_{k'=1}^K \exp(z_{k'})}, \qquad k = 1, \dots, K. \] In vector form, \[ \begin{aligned} \mathbf{y} = \left[ \begin{matrix} y_1 \\ y_2 \\ \vdots \\ y_K \end{matrix} \right] = \text{softmax}(\mathbf{z}) = \left[ \begin{matrix} \dfrac{\exp(z_1)}{\sum_{k'=1}^K \exp(z_{k'})} \\ \dfrac{\exp(z_2)}{\sum_{k'=1}^K \exp(z_{k'})} \\ \vdots \\ \dfrac{\exp(z_K)}{\sum_{k'=1}^K \exp(z_{k'})} \end{matrix} \right] \end{aligned} \]

Note that \(\mathbf{y}\) has the same dimension as \(\mathbf{z}\) and \(\mathbf{t}\). The softmax function has several desirable properties.

First, softmax outputs a probability distribution. Since \(\exp(\cdot) > 0\), each \(y_k\) is between \(0\) and \(1\), and dividing by the sum makes the \(y_k\) sum to \(1\). So \(\mathbf{y}\) satisfies both conditions for a probability distribution over the \(K\) classes.

Second, softmax preserves the ordering of the logits. Since \(\exp(\cdot)\) is strictly increasing, if class \(k\) has a larger logit than class \(k'\), class \(k\) also receives a larger probability than class \(k'\). As a result, the class with the highest logit is also the class with the highest probability, and this is the predicted class. In fact, we can skip the softmax calculation if we only need the class prediction. However, the probabilities give us more information than the class prediction alone: we can interpret each probability as a measure of the model’s confidence in predicting that class.

Third, softmax is smooth and differentiable everywhere. It can be viewed as a smooth, differentiable approximation to the argmax operation. Whereas argmax selects the highest-scoring class in a hard, non-differentiable manner, softmax assigns gradually varying probabilities to classes based on their relative scores. This is why applying the softmax function is necessary during training: it allows us to train the model using gradient descent.

Finally, softmax generalizes the sigmoid function. Try the following exercise to prove that the sigmoid is a special case of softmax for \(K = 2\).

Question: The softmax function computes the predicted probability for class \(k\) as \[ y_k = \frac{\exp(z_k)}{\sum_{k'=1}^K \exp(z_{k'})}. \] Plug in \(K = 2\) to write down \(y_1\) and \(y_2\) in terms of the logits \(z_1\) and \(z_2\). Then, show that \(y_1\) can be written in the form of the sigmoid function shown below for a single logit \(z\): \[\sigma(z) = \frac{1}{1 + e^{-z}}.\]

Answer:

With \(K = 2\), the softmax function produces two probabilities: \[ y_1 = \frac{\exp(z_1)}{\exp(z_1) + \exp(z_2)}, \qquad y_2 = \frac{\exp(z_2)}{\exp(z_1) + \exp(z_2)}. \] Dividing the numerator and the denominator of \(y_1\) by \(\exp(z_1)\) gives \[ \begin{aligned} y_1 & = \frac{1}{1 + \exp(z_2 - z_1)} \\ & = \frac{1}{1 + \exp\left(-(z_1 - z_2)\right)} \\ & = \sigma(z_1 - z_2). \end{aligned} \] So \(y_1 = \sigma(z)\) with the single logit \(z = z_1 - z_2\). Since the two probabilities sum to \(1\), \[ y_2 = 1 - y_1 = 1 - \sigma(z) = \sigma(-z). \]

3.4 The Decision Boundary

Let’s consider the decision boundary for softmax regression. Recall that the decision boundary is the set of points where the model’s prediction would change if the input changes a little bit.

Since the model predicts the class with the highest score, the prediction changes where the highest score switches from one class to another. This happens where two classes tie for the highest score. That is, a point \(\mathbf{x}\) lies on the decision boundary between classes \(i\) and \(j\) if their scores are the same, \[ z_i = z_j, \] and no other class has a larger score: \[ z_k \leq z_i \quad \text{for all } k \neq i, j. \] Both conditions are necessary. If \(z_i = z_j\) but another class \(k\) has a larger score, the model predicts class \(k\), and changing \(\mathbf{x}\) a little bit does not change the prediction.

The first condition \(z_i = z_j\) can be written as \[ \mathbf{w}_{i}^{\top} \mathbf{x} = \mathbf{w}_{j}^{\top} \mathbf{x}, \] or equivalently, \[ (\mathbf{w}_{i} - \mathbf{w}_{j})^{\top} \mathbf{x} = 0. \] This equation defines a hyperplane in the input space. The boundary between classes \(i\) and \(j\) is the part of this hyperplane where the second condition also holds, i.e., where \(i\) and \(j\) tie for the highest score. As a result, softmax regression has linear decision boundaries, which partition the input space into convex regions, one for each class.

Let’s consider a numeric example using our leaf classification problem. Suppose our model has the following weights, where each row contains the weights \(w_1\), \(w_2\), and the bias \(b\) for one class: \[ \left[ \begin{array}{lrrr} & w_1 & w_2 & b \\ \text{Oak} & -1 & 1 & 0 \\ \text{Maple} & 1 & 0 & -9 \\ \text{Birch} & 0 & -1 & 3 \end{array} \right] \]

The visualization below shows the decision regions for these weights. Try changing the weights to see how the decision boundaries move, and check that the boundaries always consist of straight line segments.

Multi-class decision regions

Figure 2: Editable class weights define linear scores for Oak, Maple, and Birch, partitioning the input space into decision regions.

Given a leaf with width \(x_1\) and height \(x_2\), the scores for the three classes are: \[ \begin{aligned} z_1 & = -x_1 + x_2 & \text{(Oak)} \\ z_2 & = -9 + x_1 & \text{(Maple)} \\ z_3 & = 3 - x_2 & \text{(Birch)} \end{aligned} \]

Let’s find the boundary between Oak and Maple. The first condition \(z_1 = z_2\) gives \[ -x_1 + x_2 = -9 + x_1, \] which simplifies to the line \[ x_2 = 2 x_1 - 9. \] Every point on this line satisfies the first condition, but only some of them are on the boundary. For example, at \((x_1, x_2) = (9, 9)\), the scores are \(z_1 = 0\), \(z_2 = 0\), and \(z_3 = -6\). Oak and Maple tie for the highest score, so this point is on the Oak–Maple boundary. If we increase the width a little bit, the model predicts Maple. If we decrease the width a little bit, the model predicts Oak.

To find all the points on this line that are on the boundary, we apply the second condition \(z_3 \leq z_1\). Substituting \(x_2 = 2 x_1 - 9\) into the scores gives \[ \begin{aligned} z_1 & = -x_1 + (2 x_1 - 9) = x_1 - 9, \\ z_3 & = 3 - (2 x_1 - 9) = 12 - 2 x_1. \end{aligned} \] So the second condition becomes \[ 12 - 2 x_1 \leq x_1 - 9, \] which simplifies to \[ x_1 \geq 7. \] So the Oak–Maple boundary is not the whole line, but only the part of the line that starts at \((7, 5)\) and extends in the direction of increasing \(x_1\).

Question: The points \((5, 1)\) and \((7, 5)\) are also on the line \(x_2 = 2 x_1 - 9\). For each point, compute the three scores and determine whether the point is on the Oak–Maple boundary. Check your answers using the visualization above.

Answer:

At \((x_1, x_2) = (5, 1)\), the scores are \(z_1 = -4\), \(z_2 = -4\), and \(z_3 = 2\). Oak and Maple tie, but Birch has the highest score. The model predicts Birch here, and changing the input a little bit does not change the prediction. So this point is not on the boundary. This agrees with the condition \(x_1 \geq 7\), since \(x_1 = 5\).

At \((x_1, x_2) = (7, 5)\), the scores are \(z_1 = z_2 = z_3 = -2\). All three classes tie, so this point is on all three boundaries. This is where the Oak–Maple, Oak–Birch, and Maple–Birch boundaries meet, and the model is equally uncertain among all three leaf types.

The Oak–Birch and Maple–Birch boundaries can be found in the same way.

Question: Find the decision boundary between Maple and Birch for the weights above. Check your answer using the visualization above.

Answer:

The first condition \(z_2 = z_3\) gives \[ -9 + x_1 = 3 - x_2, \] which simplifies to the line \[ x_2 = 12 - x_1. \] The second condition is \(z_1 \leq z_2\). Substituting \(x_2 = 12 - x_1\) into the scores gives \[ \begin{aligned} z_1 & = -x_1 + (12 - x_1) = 12 - 2 x_1, \\ z_2 & = -9 + x_1. \end{aligned} \] So the second condition becomes \[ 12 - 2 x_1 \leq -9 + x_1, \] which simplifies to \[ x_1 \geq 7. \] So the Maple–Birch boundary is the part of the line \(x_2 = 12 - x_1\) that starts at \((7, 5)\) and extends in the direction of increasing \(x_1\). In the visualization, it goes from \((7, 5)\) to \((12, 0)\).

Note that applying the softmax function does not change the decision boundary. Since softmax preserves the ordering of the logits, the class with the highest probability is always the class with the highest logit. So the predictions, and therefore the location and shape of the decision boundary, are entirely determined by the linear scores. Softmax only converts these scores into probabilities, which tell us how confident the model is in each prediction.

4 Multi-Class Classification: Training

We now turn from inference to training. The goal of training is to choose the parameters \(\mathbf{W}\) so that the predictions produced during inference match the labelled data as closely as possible.

4.1 Cross-Entropy Loss for Multi-Class Classification

We saw that softmax is a generalization of the sigmoid function from two classes to \(K\) classes. How can we generalize the cross-entropy loss from the previous chapter? Recall that the binary cross-entropy loss for classifying a single example is given below. \[ \mathcal{L}_{\text{CE}} (y,t) = -t \log y - (1 - t) \log(1 - y) \]

This loss function naturally generalizes to the multi-class setting for softmax regression. Given a prediction vector \(\mathbf{y}\) and a target vector \(\mathbf{t}\), the cross-entropy loss is defined as \[ \begin{aligned} \mathcal{L}_{\text{CE}} (\mathbf{y}, \mathbf{t}) &= -t_1 \log(y_1) - t_2 \log(y_2) - \dots - t_K \log(y_K) \\ &= -\sum_{k=1}^K t_k \log(y_k) \end{aligned} \]

Question: How many terms in the above expression are non-zero?

Answer:

Only one term is non-zero. The target vector \(\mathbf{t}\) is a one-hot vector (one \(t_k\) is \(1\) and the rest are \(0\)). As a result, the expression simplifies to the negative log-probability assigned to the true class. The loss therefore penalizes the model whenever it assigns low probability to the correct class.

Question: Derive the vectorized expression for the cross-entropy loss for \(K\) classes. \[ \begin{aligned} \mathcal{L}_{\text{CE}} (\mathbf{y}, \mathbf{t}) = -\sum_{k=1}^K t_k \log(y_k) \end{aligned} \]

Answer:

The two inputs are the target vector \(\mathbf{t} \in \mathbb{R}^{K \times 1}\) and the prediction vector \(\mathbf{y} \in \mathbb{R}^{K \times 1}\). The result should be a scalar loss. This suggests that we should use an inner (dot) product. The vectorized loss is given below:

\[ \mathcal{L}_{\text{CE}} (\mathbf{y}, \mathbf{t}) = -\mathbf{t}^{\top} \log( \mathbf{y} ) \]

4.2 Training Objective and Optimization

Given a dataset of \(N\) examples, the overall training objective is to minimize the average cross-entropy loss across the dataset. As we have seen before, we can accomplish this by using gradient descent.

The gradients of the loss with respect to the model parameters have a simple and intuitive form: \[ \nabla_{\mathbf{w}_k} \mathcal{L}_{\text{CE}} = \frac{\partial \mathcal{L}_{\text{CE}}}{\partial z_k} \nabla_{\mathbf{w}_k} z_k = (y_k - t_k) \mathbf{x} \]

As in binary logistic regression, learning proceeds by iteratively adjusting the parameters: \[ \mathbf{w}_k \leftarrow \mathbf{w}_k - \frac{\alpha}{N} \sum_{i=1}^{N} (y^{(i)}_k - t^{(i)}_k) \cdot \mathbf{x}^{(i)} \]

5 The Complete Model

We can now define the complete softmax regression model by combining all the pieces. First, we compute the logits using a linear function of the input. Then, we apply the softmax function to convert the logits into probabilities. Finally, we train the weights by minimizing the cross-entropy loss.

Definition: The softmax regression model for multi-class classification has the hypothesis form \[ \begin{array}{l} \text{Compute scores (logits):} \\[5pt] \qquad \mathbf{z} = \mathbf{W} \mathbf{x} \\[15pt] \text{Convert to probabilities:} \\[5pt] \qquad \mathbf{y} = \text{softmax}(\mathbf{z}), \quad \text{i.e.,} \quad y_k = \displaystyle\frac{\exp(z_k)}{\sum_{k'=1}^K \exp(z_{k'})} \\[15pt] \text{Minimize cross-entropy loss:} \\[5pt] \qquad \mathcal{L}_{\text{CE}}(\mathbf{y}, \mathbf{t}) = -\mathbf{t}^{\top} \log(\mathbf{y}) \end{array} \]

where

  • \(\mathbf{x} \in \mathbb{R}^{D+1}\) is the input vector for \(D\) features, which includes a dummy feature \(1\) for the bias,
  • \(\mathbf{W} \in \mathbb{R}^{K \times (D+1)}\) is the weight matrix, where the \(k\)-th row contains the weights and the bias for class \(k\),
  • \(\mathbf{z} \in \mathbb{R}^K\) is the vector of logits, one for each class,
  • \(\mathbf{y} \in \mathbb{R}^K\) is the vector of predicted probabilities for the \(K\) classes,
  • \(\log(\cdot)\) is applied element-wise,
  • \(\mathbf{t} \in \mathbb{R}^K\) is the one-hot target vector, and
  • \(\mathcal{L}_{\text{CE}}(\mathbf{y}, \mathbf{t})\) is the cross-entropy loss.

6 Summary

Softmax regression extends logistic regression from two classes to many classes. The overall recipe stays the same. We compute a linear score for each class, convert the scores into probabilities, and train the model by minimizing the cross-entropy loss with gradient descent. Each piece has a natural multi-class version: one-hot vectors for the targets, the softmax function in place of the sigmoid, and a cross-entropy loss that sums over all the classes.

Having more than two classes brings a few new challenges. We need a way to encode the targets that does not suggest the classes have an order, which is why we use one-hot vectors instead of integers. The model needs a separate set of weights for every class, so it grows as we add more classes. The decision boundaries are also harder to reason about. Instead of one line separating two classes, we have boundaries between pairs of classes, and each boundary only covers the part of the space where those two classes tie for the highest score.

This chapter illustrates Fundamental Idea #4 (ML describes geometric processes). The model divides the input space into regions, one for each class, and the boundaries between these regions are linear. Like logistic regression, softmax regression is simple and easy to interpret, but it cannot represent non-linear decision boundaries. Applying the softmax function does not change the decision regions. Instead, it reflects Fundamental Idea #5 (ML demands a probabilistic lens). The model outputs a probability for every class, which tells us not only which class it predicts but also how confident it is.