Multivariate Gaussian Discriminant Analysis

1 Introduction

In the previous chapter we reviewed the multivariate Gaussian distribution and how to learn its parameters from data. We now use that distribution as the class-conditional model for multivariate Gaussian Discriminant Analysis (GDA). We assume that the feature vector \(\mathbf{x} \in \mathbb{R}^D\) given the class \(c\) is drawn from a multivariate Gaussian. Each class \(k\) has its own mean vector \(\boldsymbol{\mu}_k\) and covariance matrix \(\boldsymbol{\Sigma}_k\). We fit these parameters from data (e.g., by maximum likelihood), then classify new points using Bayes’ rule.

To illustrate this idea, we use the same leaf dataset as before, but now with two features leaf width (\(x_1\)) and leaf height (\(x_2\)).

Table 1: Leaf prediction training data (same as in the supervised learning chapter).
# Leaf Width Leaf Height Leaf Type
1 7.0 cm 12.0 cm Oak
2 9.0 cm 6.0 cm Oak
3 10.0 cm 18.0 cm Oak
4 5.0 cm 10.0 cm Oak
5 13.0 cm 11.0 cm Maple
6 11.0 cm 9.0 cm Maple
7 9.0 cm 14.0 cm Maple

Hopefully, you will find that learning a multivariate GDA model for this classification problem mirrors what we did with the univariate GDA model. The only difference is that \(p(\mathbf{x}|c)\) is a multivariate Gaussian instead of a univariate one.

One difference, however, is that a multivariate Gaussian has many more learnable parameters than a univariate Gaussian. This increase in model capacity can be problematic (prone to overfitting). We will explore strategies for constraining the covariance matrix. For example, the covariance matrix can be restricted to be diagonal for Gaussian Naive Bayes, or shared across classes. These reduce parameters and also result in changes to the shape of the decision boundary.

2 The Multivariate GDA Model

As in a univariate GDA model, a (multivariate) GDA model is a generative model of the joint distribution of the feature and the class. We again write \[\begin{align*} p(\mathbf{x}, c) = p(\mathbf{x} \mid c)\,p(c). \end{align*}\] The “Gaussian” part of GDA now means that we assume that given the class \(c\), the feature \(\mathbf{x}\) is drawn from a multivariate Gaussian distribution. So for each class \(k \in \{0, 1\}\) we have \[\begin{align*} p(\mathbf{x} \mid c = k) = \mathcal{N}(\mathbf{x}; \boldsymbol{\mu}_k, \boldsymbol{\Sigma}_k) = \frac{1}{(2\pi)^{D/2}\,|\boldsymbol{\Sigma}_k|^{1/2}} \exp\left(-\frac{1}{2}(\mathbf{x} - \boldsymbol{\mu}_k)^\top \boldsymbol{\Sigma}_k^{-1} (\mathbf{x} - \boldsymbol{\mu}_k)\right), \end{align*}\] with class-specific mean vector \(\boldsymbol{\mu}_k\) and covariance matrix \(\boldsymbol{\Sigma}_k\).

Assuming a binary classification problem, the prior over the class is again \(\mathrm{Bernoulli}(\theta)\): \(p(c=1) = \theta\), \(p(c=0) = 1 - \theta\).

Question: Multivariate GDA as described above assumes binary classes (\(c \in \{0,1\}\)) with a Bernoulli prior \(p(c)\). How would you extend the model to \(K > 2\) classes?

Answer:

Replace the Bernoulli prior with a Categorical distribution over \(K\) classes: \(p(c = k) = \theta_k\) for \(k = 0, \ldots, K-1\), with \(\sum_k \theta_k = 1\). Each class still has its own \(\boldsymbol{\mu}_k\) and \(\boldsymbol{\Sigma}_k\), learned from the class-\(k\) training points. The MLE for each \(\theta_k\) is the fraction of training points in class \(k\).

As a running example, we use the same leaf dataset as before. Again, let’s use the class label \(c \in \{0, 1\}\) (e.g., \(c=0\) for oak, \(c=1\) for maple). Our observations are the training set \(\mathcal{D} = \{(\mathbf{x}^{(1)}, c^{(1)}), (\mathbf{x}^{(2)}, c^{(2)}), \dots, (\mathbf{x}^{(N)}, c^{(N)})\}\).

Question: List all the parameters of the binary multivariate GDA model. How many scalars must be estimated when \(D = 2\)?

Answer:

The parameters are \(\boldsymbol{\mu}_0\), \(\boldsymbol{\Sigma}_0\), \(\boldsymbol{\mu}_1\), \(\boldsymbol{\Sigma}_1\), and \(\theta\).

For \(D=2\): each mean vector has \(2\) entries; each covariance matrix has \(D(D+1)/2 = 3\) free entries (symmetric). Total: \(2 + 3 + 2 + 3 + 1 = \mathbf{11}\) parameters.

Together these parameters fully specify the model. Once learned, we classify new \(\mathbf{x}\) by computing the posterior \(p(c \mid \mathbf{x})\) via Bayes’ rule and predicting the class with the highest posterior probability.

2.1 Learning Multivariate GDA

The parameters are fit by maximum likelihood, following the same derivation as in the previous chapter. For each class \(k\) we use only the training points from that class; the prior is the empirical class fraction. The MLE formulas mirror the univariate case, but now we use vectors and matrices.

Let \(N_k\) be the number of training points in class \(k\), and let \(r_k^{(i)} = \mathbb{I}(c^{(i)} = k)\) be indicator variables for whether data point \(i\) belongs to class \(k\). Then our MLE estimates of the parameters are:

\[\begin{align*} \hat{\theta} &= \frac{N_1}{N}, \\ \hat{\boldsymbol{\mu}}_k &= \frac{1}{N_k}\sum_{i : c^{(i)}=k} \mathbf{x}^{(i)} = \frac{\sum_{i=1}^N r_k^{(i)} \mathbf{x}^{(i)}}{\sum_{i=1}^N r_k^{(i)}}, \\ \hat{\boldsymbol{\Sigma}}_k &= \frac{1}{N_k}\sum_{i : c^{(i)}=k} (\mathbf{x}^{(i)} - \hat{\boldsymbol{\mu}}_k)(\mathbf{x}^{(i)} - \hat{\boldsymbol{\mu}}_k)^\top. \end{align*}\]

So for each class we compute the sample mean and sample covariance using only the data from that class; \(\theta\) is the fraction of training points in class 1.

For the leaf dataset, the fitting of the Gaussians was shown in the previous chapter. We had, for oak (\(k=0\)) and maple (\(k=1\)):

\[\begin{align*} \hat{\boldsymbol{\mu}}_0 &= \begin{bmatrix}7.75 \\ 11.5\end{bmatrix}\\ \quad\hat{\boldsymbol{\Sigma}}_0 &= \begin{bmatrix}3.69 & 2.88\\2.88 & 18.75\end{bmatrix} \\ \hat{\boldsymbol{\mu}}_1 &= \begin{bmatrix}11 \\ 11.3\end{bmatrix}\\ \quad\hat{\boldsymbol{\Sigma}}_1 &= \begin{bmatrix}2.67 & -2.00\\-2.00 & 4.22\end{bmatrix} \end{align*}\]

With \(N=7\), \(N_0=4\) (oak) and \(N_1=3\) (maple), the prior MLE is \(\hat{\theta} = N_1/N = 3/7 \approx 0.43\).

Fitted class-conditional Gaussians

Figure 1: Class Oak (left) and Maple (right).

2.2 Inference

At test time we observe a new feature vector \(\mathbf{x}\) and wish to predict the class \(c\). As in the univariate case, we use Bayes’ rule: \[ p(c \mid \mathbf{x}) = \frac{p(\mathbf{x} \mid c)\,p(c)}{\sum_{c'} p(\mathbf{x} \mid c')\,p(c')}. \] We plug in the learned multivariate Gaussians \(p(\mathbf{x} \mid c=0)\) and \(p(\mathbf{x} \mid c=1)\) and the prior \(p(c)\). We then predict the class with the higher posterior probability; for binary classification, we predict class 1 (maple) if \(p(c=1 \mid \mathbf{x}) \geq 1/2\) and class 0 (oak) otherwise.

In practice it is often convenient to work with log-posteriors. Ignoring terms that do not depend on the class, \[ \log p(c=k \mid \mathbf{x}) \propto -\frac{1}{2}\log|\boldsymbol{\Sigma}_k| - \frac{1}{2}(\mathbf{x} - \boldsymbol{\mu}_k)^\top \boldsymbol{\Sigma}_k^{-1}(\mathbf{x} - \boldsymbol{\mu}_k) + \log p(c=k). \] We can compute this for \(k=0\) and \(k=1\) and choose the class that gives the larger value.

Question: Using the fitted GDA model, classify a new leaf with \[\mathbf{x} = \begin{bmatrix}12 \\ 16\end{bmatrix}\] Compute the log-posterior score for each class and state the predicted label.

Answer: Let’s use the symbol \(z_k\) to represent the scores for each class \(k=0,1\): \[z_k = -\frac{1}{2}\log|\hat{\boldsymbol{\Sigma}}_k| - \frac{1}{2}(\mathbf{x}-\hat{\boldsymbol{\mu}}_k)^\top\hat{\boldsymbol{\Sigma}}_k^{-1}(\mathbf{x}-\hat{\boldsymbol{\mu}}_k) + \log\hat{\theta}_k\]

For oak (\(k=0\)), we have \[\begin{align*} |\hat{\boldsymbol{\Sigma}}_0| &= 60.9 \\ \hat{\boldsymbol{\Sigma}}_0^{-1} &= \frac{1}{60.9}\begin{bmatrix}18.75 & -2.88\\-2.88 & 3.69\end{bmatrix} \\ (\mathbf{x}-\hat{\boldsymbol{\mu}}_0) &= \begin{bmatrix}4.25\\4.5\end{bmatrix} \\ (\mathbf{x}-\hat{\boldsymbol{\mu}}_0)^\top\hat{\boldsymbol{\Sigma}}_0^{-1}(\mathbf{x}-\hat{\boldsymbol{\mu}}_0) &\approx 4.98 \\ z_0 &= -\tfrac{1}{2}(4.11) - \tfrac{1}{2}(4.98) + (-0.56) \approx -5.1. \end{align*}\]

For maple (\(k=1\)), we have \[\begin{align*} |\hat{\boldsymbol{\Sigma}}_1| &= 7.3 \\ \hat{\boldsymbol{\Sigma}}_1^{-1} &= \frac{1}{7.3}\begin{bmatrix}4.22 & 2\\2 & 2.67\end{bmatrix} \\ (\mathbf{x}-\hat{\boldsymbol{\mu}}_1) &= \begin{bmatrix}12-11\\16-34/3\end{bmatrix} = \begin{bmatrix}1\\14/3\end{bmatrix} \\ (\mathbf{x}-\hat{\boldsymbol{\mu}}_1)^\top\hat{\boldsymbol{\Sigma}}_1^{-1}(\mathbf{x}-\hat{\boldsymbol{\mu}}_1) &\approx 11.1 \\ z_1 &= -\tfrac{1}{2}(1.99) - \tfrac{1}{2}(11.1) + (-0.85) \approx -7.4. \end{align*}\]

Since \(z_0 > z_1\), the model’s prediction is \(c=0\), or oak.

If we require a posterior probability, we can compute \[ p\!\left(c=0 \mid \mathbf{x}\right) = \frac{e^{z_0}}{e^{z_0}+e^{z_1}} \approx 0.91. \] So the model assigns about 91% probability to oak and 9% to maple.

2.3 Decision Boundary

The decision boundary is the set of \(\mathbf{x}\) where \(p(c=0 \mid \mathbf{x}) = p(c=1 \mid \mathbf{x})\), i.e., where the two (weighted) class-conditional densities are equal. Setting the log-posteriors equal and rearranging yields \[ (\mathbf{x} - \boldsymbol{\mu}_0)^\top \boldsymbol{\Sigma}_0^{-1}(\mathbf{x} - \boldsymbol{\mu}_0) = (\mathbf{x} - \boldsymbol{\mu}_1)^\top \boldsymbol{\Sigma}_1^{-1}(\mathbf{x} - \boldsymbol{\mu}_1) + \text{Const}. \]

For the leaf dataset the boundary is a quadratic curve in \((x_1, x_2)\), since the two classes have different covariance matrices. The figure below shades the oak region (blue) and maple region (orange) and overlays the class contours and training data.

GDA decision regions

Figure 2: Shading shows which class has higher posterior probability at each point. Ellipses are the innermost two contour levels of each fitted Gaussian. Training points use the same symbols as throughout this chapter.

This is a quadratic equation in \(\mathbf{x}\), so the decision boundary is a conic section (ellipse, parabola, or hyperbola). In 2D, we typically get a curved boundary that separates the two class regions.

The figure below lets you explore how the parameters of each class Gaussian—and the prior \(\theta\)—jointly shape the decision boundary. The shaded regions show which class has the higher posterior at each point; the ellipses are the two innermost contour levels.

Interactive GDA decision boundary

Figure 3: Sliders control \(\boldsymbol{\mu}_k\), \(\boldsymbol{\Sigma}_k\), and \(\theta\) for each class. The decision region updates in real time.

Question: Use the interactive figure above to produce each of the following boundary shapes: linear boundary, elliptical boundary enclosing one class, hyperbolic boundary.

3 Variants: Restrictions on Covariance

GDA is an expressive model. One challenge with GDA is that the number of parameters that we need to fit grows quadratically with respect to the number of features \(D\). By default, GDA captures the pairwise correlations between each pair of features. As we saw in the Naive Bayes section, this is not scalable when \(D\) is large.

To get around this issue and reduce the number of learnable parameters, we can place restrictions on the “shape” of the covariance matrix based on what we know about our data. The following variants of GDA do exactly that.

3.1 Linear Discriminant Analysis

In this variant, we restrict the covariance matrix of the different classes to be identical. The single covariance matrix can be learned via MLE using \[\begin{align*} \hat{\boldsymbol{\Sigma}} = \frac{1}{N}\sum_{i=1}^{N}\bigl(\mathbf{x}^{(i)} - \hat{\boldsymbol{\mu}}_{c^{(i)}}\bigr)\bigl(\mathbf{x}^{(i)} - \hat{\boldsymbol{\mu}}_{c^{(i)}}\bigr)^\top, \end{align*}\]

where \(\hat{\boldsymbol{\mu}}_{c^{(i)}}\) denotes the estimated class mean for the class of point \(i\) (i.e., \(\hat{\boldsymbol{\mu}}_0\) if \(c^{(i)}=0\) and \(\hat{\boldsymbol{\mu}}_1\) if \(c^{(i)}=1\)). This variant of GDA is called Linear Discriminant Analysis.

Definition: Linear Discriminant Analysis (LDA) is a special case of GDA in which all classes share the same covariance matrix \(\boldsymbol{\Sigma}\). Because the quadratic terms in the log-posterior cancel, the resulting decision boundary is a hyperplane (linear in \(\mathbf{x}\)), hence the name.

For the leaf classification problem, applying this constraint gives a pooled covariance \(\hat{\boldsymbol{\Sigma}} \approx \begin{bmatrix}3.25 & 0.79\\0.79 & 12.52\end{bmatrix}\). The resulting decision boundary is shown in Figure 4 below.

LDA decision boundary

Figure 4: The boundary is a straight line.

Notice that the decision boundary in Figure 4 is linear. This linearity is not an accident, and can be demonstrated mathematically. The decision boundary is where \[-\tfrac{1}{2}\log|\boldsymbol{\Sigma}_0| - \tfrac{1}{2}(\mathbf{x}-\boldsymbol{\mu}_0)^\top\boldsymbol{\Sigma}_0^{-1}(\mathbf{x}-\boldsymbol{\mu}_0) + \log(1-\theta) = -\tfrac{1}{2}\log|\boldsymbol{\Sigma}_1| - \tfrac{1}{2}(\mathbf{x}-\boldsymbol{\mu}_1)^\top\boldsymbol{\Sigma}_1^{-1}(\mathbf{x}-\boldsymbol{\mu}_1) + \log\theta.\] When \(\boldsymbol{\Sigma}_0 = \boldsymbol{\Sigma}_1 = \boldsymbol{\Sigma}\), the \(\log|\boldsymbol{\Sigma}|\) terms cancel. Expanding each quadratic using \[(\mathbf{x}-\boldsymbol{\mu}_k)^\top\boldsymbol{\Sigma}^{-1}(\mathbf{x}-\boldsymbol{\mu}_k) = \mathbf{x}^\top\boldsymbol{\Sigma}^{-1}\mathbf{x} - 2\boldsymbol{\mu}_k^\top\boldsymbol{\Sigma}^{-1}\mathbf{x} + \boldsymbol{\mu}_k^\top\boldsymbol{\Sigma}^{-1}\boldsymbol{\mu}_k,\] the \(\mathbf{x}^\top\boldsymbol{\Sigma}^{-1}\mathbf{x}\) terms also cancel, leaving \[(\boldsymbol{\mu}_1-\boldsymbol{\mu}_0)^\top\boldsymbol{\Sigma}^{-1}\mathbf{x} = \text{const},\] which is linear in \(\mathbf{x}\). The decision boundary is therefore a hyperplane.

Use the figure below to verify this: no matter how you adjust the means, the shared covariance entries, or the prior, the boundary remains a straight line.

Interactive shared-covariance GDA

Figure 5: Both classes share one \(\boldsymbol{\Sigma}\); the \(\sigma_{12}\) knob rotates the shared ellipses but the boundary stays linear.

Question: For the leaf dataset with \(D=2\) features, how many parameters does LDA have? Compare to full GDA.

Answer:

LDA has two class means (\(2 \times D = 4\) entries), one shared covariance matrix (\(D(D+1)/2 = 3\) free entries for \(D=2\)), and the prior \(\theta\): \[2 + 2 + 3 + 1 = 8 \text{ parameters,}\] compared to 11 for full GDA. The saving comes from using a single shared \(\boldsymbol{\Sigma}\) instead of two separate ones.

In general, binary LDA with \(D\) features has \(2D + D(D+1)/2 + 1\) parameters—two means of \(D\) entries each, one shared symmetric covariance with \(D(D+1)/2\) free entries, and \(\theta\). With fewer parameters than full GDA, LDA is less prone to overfitting.

3.2 Gaussian Naive Bayes

In this variant, we assume that the features are conditionally independent given the class, so \(p(\mathbf{x} \mid c = k) = \prod_{j=1}^D p(x_j \mid c=k)\). Modeling each marginal as a univariate Gaussian is equivalent to multivariate GDA with a diagonal covariance matrix: all off-diagonal (correlation) entries are forced to zero. The MLE for each per-feature variance is the usual sample variance within each class: \[\begin{align*} \hat{\sigma}_{kj}^2 = \frac{1}{N_k}\sum_{i:\,c^{(i)}=k}(x_j^{(i)} - \hat{\mu}_{kj})^2, \end{align*}\] so \(\hat{\boldsymbol{\Sigma}}_k = \operatorname{diag}(\hat{\sigma}_{k1}^2,\ldots,\hat{\sigma}_{kD}^2)\).

Definition: Gaussian Naive Bayes (GNB) is GDA with a diagonal covariance matrix per class: \(\boldsymbol{\Sigma}_k = \operatorname{diag}(\sigma_{k1}^2, \ldots, \sigma_{kD}^2)\).

Question: Show that in a GNB model, the conditional independence assumption \(p(\mathbf{x} \mid c=k) = \prod_{j=1}^D p(x_j \mid c=k)\) holds.

Answer:

With a diagonal covariance \(\boldsymbol{\Sigma}_k = \operatorname{diag}(\sigma_{k1}^2,\ldots,\sigma_{kD}^2)\), the determinant factors as \(|\boldsymbol{\Sigma}_k| = \prod_j \sigma_{kj}^2\) and the inverse is \(\boldsymbol{\Sigma}_k^{-1} = \operatorname{diag}(1/\sigma_{k1}^2, \ldots, 1/\sigma_{kD}^2)\). The quadratic form therefore decomposes feature-by-feature: \[(\mathbf{x}-\boldsymbol{\mu}_k)^\top\boldsymbol{\Sigma}_k^{-1}(\mathbf{x}-\boldsymbol{\mu}_k) = \sum_{j=1}^D \frac{(x_j - \mu_{kj})^2}{\sigma_{kj}^2}.\] Substituting into the multivariate Gaussian density: \[p(\mathbf{x}\mid c=k) = \frac{1}{(2\pi)^{D/2}\prod_j\sigma_{kj}}\exp\!\left(-\frac{1}{2}\sum_{j=1}^D\frac{(x_j-\mu_{kj})^2}{\sigma_{kj}^2}\right) = \prod_{j=1}^D \frac{1}{\sqrt{2\pi}\,\sigma_{kj}}\exp\!\left(-\frac{(x_j-\mu_{kj})^2}{2\sigma_{kj}^2}\right),\] which is exactly \(\prod_{j=1}^D p(x_j \mid c=k)\) where each factor is a univariate Gaussian.

For the leaf dataset, fitting diagonal covariances gives \[\hat{\boldsymbol{\Sigma}}_0 = \begin{bmatrix}3.69 & 0\\0 & 18.75\end{bmatrix}, \qquad \hat{\boldsymbol{\Sigma}}_1 = \begin{bmatrix}2.67 & 0\\0 & 4.22\end{bmatrix}.\] The contour ellipses are now axis-aligned, as shown in Figure 6.

GNB decision boundary

Figure 6: Contour ellipses are axis-aligned because off-diagonal covariance entries are zeroed out.

Unlike LDA, the boundary in Figure 6 is curved, not a straight line. However, unlike full GDA, the boundary is an axis-aligned conic section (no tilting), because the off-diagonal entries are zero. Again, the decision boundary is where \[-\tfrac{1}{2}\log|\boldsymbol{\Sigma}_0| - \tfrac{1}{2}(\mathbf{x}-\boldsymbol{\mu}_0)^\top\boldsymbol{\Sigma}_0^{-1}(\mathbf{x}-\boldsymbol{\mu}_0) + \log(1-\theta) = -\tfrac{1}{2}\log|\boldsymbol{\Sigma}_1| - \tfrac{1}{2}(\mathbf{x}-\boldsymbol{\mu}_1)^\top\boldsymbol{\Sigma}_1^{-1}(\mathbf{x}-\boldsymbol{\mu}_1) + \log\theta.\] Because \(\boldsymbol{\Sigma}_k^{-1}\) is diagonal, the quadratic form reduces to a sum of per-feature terms with no cross-products: \[(\mathbf{x}-\boldsymbol{\mu}_k)^\top\boldsymbol{\Sigma}_k^{-1}(\mathbf{x}-\boldsymbol{\mu}_k) = \sum_{j=1}^D \frac{(x_j - \mu_{kj})^2}{\sigma_{kj}^2}.\] Substituting and rearranging, we have the decision boundary \[\sum_{j=1}^D \left(\frac{1}{2\sigma_{1j}^2} - \frac{1}{2\sigma_{0j}^2}\right)x_j^2 - \sum_{j=1}^D\left(\frac{\mu_{1j}}{\sigma_{1j}^2} - \frac{\mu_{0j}}{\sigma_{0j}^2}\right)x_j = \text{const.}\] This is quadratic in each \(x_j\) individually. Crucially, there are no cross terms \(x_j x_l\) with \(j \neq l\), because the inverse of a diagonal matrix is also diagonal. When \(\sigma_{0j}^2 \neq \sigma_{1j}^2\), the \(x_j^2\) coefficients are non-zero and the boundary is a quadratic curve; but since only squared individual features appear, the conic sections are always axis-aligned.

The figure below shows the decision boundary for GNB for D=2. Use this to explore the possible decision boundaries.

Interactive Gaussian Naive Bayes

Figure 7: Each class has two independent variance parameters; there is no \(\sigma_{12}\) knob because off-diagonal covariances are always zero.

Question: For the leaf dataset with \(D=2\) features, how many parameters does GNB have? Compare to full GDA and LDA.

Answer:

GNB has two class means (\(2 \times D = 4\) entries), two diagonal covariance matrices (\(D\) entries each, so \(2D = 4\) total), and the prior \(\theta\): \[2 + 2 + 2 + 2 + 1 = 9 \text{ parameters,}\] compared to 11 for full GDA and 8 for LDA. The saving over full GDA comes from zeroing the off-diagonal entries; LDA saves further by sharing the covariance across classes.

In general, binary GNB with \(D\) features has \(4D + 1\) parameters—two means of \(D\) entries each, two diagonal covariances of \(D\) entries each, and \(\theta\). This is a marked reduction in the number of parameters, and the number of parameters is now linear in \(D\).

3.3 Isotropic Covariance

A further simplification constrains each class covariance to a scalar multiple of the identity: \(\boldsymbol{\Sigma}_k = \sigma_k^2 \mathbf{I}\). This means every feature has the same variance within a class and all features are uncorrelated. The contours become circles rather than ellipses, and classification reduces to a distance comparison weighted by the class-specific spread. The MLE scalar variance for class \(k\) is the average of the per-feature sample variances: \[\hat{\sigma}_k^2 = \frac{1}{D}\sum_{j=1}^D \hat{\sigma}_{kj}^2.\]

Definition: Isotropic GDA is GDA with \(\boldsymbol{\Sigma}_k = \sigma_k^2 \mathbf{I}\): a single scalar variance per class, equal in all directions. Each class is modeled by a spherical Gaussian.

For the leaf dataset, the isotropic variances are \(\hat{\sigma}_0^2 \approx 11.22\) and \(\hat{\sigma}_1^2 \approx 3.44\). Because oak leaves vary considerably more in size than maple leaves, the two circles are very different in scale, which pushes the decision boundary away from the larger class.

Isotropic GDA decision boundary

Figure 8: Contours are circles; the large difference in spread (\(\hat{\sigma}_0^2 \approx 11.22\) vs \(\hat{\sigma}_1^2 \approx 3.44\)) produces a curved boundary.

The decision boundary is again where \(z_0 = z_1\). With \(\boldsymbol{\Sigma}_k = \sigma_k^2 \mathbf{I}\), the log-posterior simplifies to \[z_k = -D\log\sigma_k - \frac{\|\mathbf{x} - \boldsymbol{\mu}_k\|^2}{2\sigma_k^2} + \log p(c=k).\] Setting \(z_0 = z_1\), the \(\|\mathbf{x}\|^2\) terms appear with coefficient \(\tfrac{1}{2\sigma_1^2} - \tfrac{1}{2\sigma_0^2}\). When \(\sigma_0 = \sigma_1\), this coefficient is zero and the \(\|\mathbf{x}\|^2\) terms cancel, leaving an expression linear in \(\mathbf{x}\)—exactly like LDA. When \(\sigma_0 \neq \sigma_1\), the quadratic terms remain and the boundary is a circle (or more generally a conic section).

Use the figure below to verify: set both variance sliders equal and the boundary becomes a straight line; as soon as they differ it curves again.

Interactive isotropic GDA

Figure 9: Each class has a single scalar variance \(\sigma_k^2\); set both sliders equal to see the boundary become linear.

Question: For the leaf dataset with \(D=2\) features, how many parameters does isotropic GDA have? Compare to full GDA, LDA, and GNB.

Answer:

Isotropic GDA has two class means (\(2 \times D = 4\) entries), two scalar variances (one per class), and the prior \(\theta\): \[4 + 2 + 1 = 7 \text{ parameters,}\] compared to 11 for full GDA, 8 for LDA, and 9 for GNB. Replacing the full covariance matrix with a single number per class gives the most compact model of the four.

In general, binary isotropic GDA with \(D\) features has \(2D + 3\) parameters—two means of \(D\) entries each, two scalar variances, and \(\theta\).

4 Summary

Multivariate GDA is a generative classifier that models each class’s data as a multivariate Gaussian. Just like in univariate GDA, parameters can be learned using maximum likelihood estimation (MLE). Each class’s distribution can be learned separately from the other classes, giving closed-form estimates that are fast to fit and interpretable. Inference applies Bayes’ rule to compute the posterior \(p(c \mid \mathbf{x})\).

The covariance matrix is both the richest and most expensive part of the model. Imposing structure on it trades capacity for parsimony: sharing a single \(\boldsymbol{\Sigma}\) across classes linearises the boundary and yields Linear Discriminant Analysis, constraining \(\boldsymbol{\Sigma}_k\) to be diagonal gives Gaussian Naive Bayes, and constraining it to \(\sigma_k^2 \mathbf{I}\) gives spherical class-conditional Gaussians with scores based on distance to the class mean, scaled by variance and adjusted by priors. Each simplification reduces the number of parameters and can improve generalisation when data are scarce relative to \(D\).

This chapter illustrates both Fundamental Idea #4 (Geometric Processes) and Fundamental Idea #5 (Probabilistic Lens). The covariance matrix encodes the geometry of each class’s distribution, its orientation and spread, and the spectral decomposition makes that geometry precise. At the same time, the entire model is probabilistic: we specify a data-generating process, learn it by maximising likelihood, and classify by Bayes’ rule. The decision boundaries in the figures above are not hand-drawn separators but the natural consequence of fitting probability distributions to data.