Multivariate Gaussian Distribution
1 Introduction
This chapter introduces the multivariate Gaussian (multivariate normal) distribution and how to learn its parameters from data—material we need for multivariate Gaussian Discriminant Analysis (GDA) in the next chapter. We define the distribution, show how to fit a single Gaussian from data via maximum likelihood, and interpret the covariance matrix (including eigenvectors and contours).
2 Visualizing a Multivariate Gaussian
The multivariate Gaussian distribution is defined over a \(D\)-dimensional random vector \(\mathbf{x} \in \mathbb{R}^D\), parameterized by a mean vector \(\boldsymbol{\mu} \in \mathbb{R}^D\) and a covariance matrix \(\boldsymbol{\Sigma}\)—a \(D \times D\) symmetric, positive semi-definite matrix. The probability density function is
\[\begin{align*} p(\mathbf{x}; \boldsymbol{\mu}, \boldsymbol{\Sigma}) &= \mathcal{N}(\mathbf{x}; \boldsymbol{\mu}, \boldsymbol{\Sigma}) \\ &= \frac{1}{(2\pi)^{D/2}\,|\boldsymbol{\Sigma}|^{1/2}} \exp\left(-\frac{1}{2}(\mathbf{x} - \boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x} - \boldsymbol{\mu})\right), \end{align*}\]
where \(|\boldsymbol{\Sigma}|\) denotes the determinant of \(\boldsymbol{\Sigma}\). When \(D=1\), the multivariate formula reduces to the univariate Gaussian: \(\boldsymbol{\mu}\) is the scalar mean \(\mu\), and \(\boldsymbol{\Sigma}\) is the scalar variance \(\sigma^2\), so \(|\boldsymbol{\Sigma}|^{1/2} = \sigma\). Substituting into the PDF above yields \(\mathcal{N}(x; \mu, \sigma^2) = \frac{1}{\sqrt{2\pi}\,\sigma}\exp\bigl(-\frac{(x-\mu)^2}{2\sigma^2}\bigr)\), which is the usual univariate Gaussian density.
The mean vector and covariance matrix can also be written \[ \boldsymbol{\mu} = \begin{bmatrix} \mu_1 \\ \vdots \\ \mu_D \end{bmatrix} \quad\quad \boldsymbol{\Sigma} = \begin{bmatrix} \sigma_1^2 & \sigma_{12} & \cdots & \sigma_{1D} \\ \sigma_{12} & \sigma_2^2 & \cdots & \sigma_{2D} \\ \vdots & \vdots & \ddots & \vdots \\ \sigma_{D1} & \sigma_{D2} & \cdots & \sigma_D^2 \end{bmatrix}. \] These entries have similar interpretations as in the univariate case:
- \(\mu_j\) is the mean of the \(j\)-th variable (\(x_j\))
- \(\sigma_j^2\) is the variance of the \(j\)-th variable (\(x_j\)), and
- for \(j \neq \ell\), \(\sigma_{j\ell}\) is the covariance between the \(j\)-th and \(\ell\)-th variables (a measure of their linear association).
The matrix \(\boldsymbol{\Sigma}\) is symmetric and positive semi-definite. Together, \(\boldsymbol{\mu}\) and \(\boldsymbol{\Sigma}\) uniquely define the multivariate Gaussian.
Definition: A matrix \(\boldsymbol{\Sigma}\) is symmetric if \(\sigma_{j\ell} = \sigma_{\ell j}\) for all \(j, \ell\). Symmetric matrices have \(\boldsymbol{\Sigma} = \boldsymbol{\Sigma}^\top\). Covariance matrices are always symmetric because the covariance between \(x_j\) and \(x_\ell\) equals the covariance between \(x_\ell\) and \(x_j\).
Definition: A symmetric matrix \(\boldsymbol{\Sigma}\) is positive semi-definite (PSD) if \(\mathbf{v}^\top \boldsymbol{\Sigma} \mathbf{v} \geq 0\) for every vector \(\mathbf{v}\). Geometrically, this means \(\boldsymbol{\Sigma}\) never “flips” a direction, so all its eigenvalues are non-negative.
When \(D=2\), we can visualize the probability density function of a Gaussian using a contour plot. Here, the two axes range over the values of the features \(x_1\) and \(x_2\). The contour helps show the value of \(p({\bf x})\): recall that contours shown are curves along which \(p({\bf x})\) is constant.
The figure below illustrates how \(\boldsymbol{\mu}\) and \(\boldsymbol{\Sigma}\) affect the contour shape.
Shifting the mean
Scaling diagonal covariance
Unequal diagonal covariance
Off-diagonal covariance
Here, the probability density \(p({\bf x})\) is the largest at \(\boldsymbol{\mu}\). Thus, modifying \(\boldsymbol{\mu}\) shifts the distribution –just like in the univariate case. Modifying \(\sigma_1\) or \(\sigma_2\) changes the spread of the distribution along the corresponding axis. Finally, \(\sigma_{12}\) changes the correlation structure between the two features: when \(\sigma_{12}=0\), the two features are uncorrelated. but when \(\sigma_{12}\) is nonzero, information about one feature provides information about the other.
The figure below lets you explore how \(\boldsymbol{\mu}\) and \(\boldsymbol{\Sigma}\) affect the contour shape. Use the sliders to change the mean (\(\mu_0\), \(\mu_1\)), the diagonal entries of \(\boldsymbol{\Sigma}\) (\(\sigma_1^2\), \(\sigma_2^2\)), and the off-diagonal covariance \(\sigma_{12}\).
Interactive bivariate Gaussian
3 Learning a Multivariate Gaussian
As before, a probabilistic distribution is often a useful summary of a dataset and its underlying phenomenon. As in the univariate case, learning a multivariate Gaussian distribution to model a phenomenon provides a compact summary of the data.
We will continue with the running example from the leaf dataset, and fit a multivariate Gaussian distribution. We again take only the oak leaves from the leaf dataset and use both width and height: \(\mathbf{x}^{(1)} = (7, 12)^\top\), \(\mathbf{x}^{(2)} = (9, 6)^\top\), \(\mathbf{x}^{(3)} = (10, 18)^\top\), \(\mathbf{x}^{(4)} = (5, 10)^\top\) (all in cm), so \(N=4\) and \(D=2\). We will assume each \(\mathbf{x}^{(i)}\) is drawn from \(\mathcal{N}(\mathbf{x}; \boldsymbol{\mu}, \boldsymbol{\Sigma})\).
Question: How many parameters must be estimated to fully specify a \(D\)-dimensional multivariate Gaussian \(\mathcal{N}(\boldsymbol{\mu}, \boldsymbol{\Sigma})\)? What is the count when \(D=2\) (e.g., for the bivariate oak-leaf Gaussian)?
Answer:
The mean vector \(\boldsymbol{\mu}\) has \(D\) entries. The covariance matrix \(\boldsymbol{\Sigma}\) is \(D \times D\) and symmetric, so it has \(1 + 2 + \cdots + D = D(D+1)/2\) free parameters (diagonal plus upper triangle). In total: \(D + D(D+1)/2 = D(D+3)/2\) parameters.
For \(D=2\): mean has 2 parameters; \(\boldsymbol{\Sigma}\) has \(2(2+1)/2 = 3\) parameters (e.g., \(\sigma_1^2\), \(\sigma_2^2\), \(\sigma_{12}\)). Total: \(5\) parameters.
We again use the maximum likelihood estimation (MLE) criterion. We will choose the values of the parameters \(\boldsymbol{\mu}, \boldsymbol{\Sigma}\) that make the observed data most likely.
Assume the samples \(\mathbf{x}^{(1)}, \ldots, \mathbf{x}^{(N)}\) are drawn independently from \(\mathcal{N}(\boldsymbol{\mu}, \boldsymbol{\Sigma})\). The likelihood of the parameters given the data is the probability of the data under the model:
\[\begin{align*} \mathcal{L}(\boldsymbol{\mu}, \boldsymbol{\Sigma}) = \prod_{i=1}^{N} \mathcal{N}(\mathbf{x}^{(i)}; \boldsymbol{\mu}, \boldsymbol{\Sigma}) = \prod_{i=1}^{N} \frac{1}{(2\pi)^{D/2} |\boldsymbol{\Sigma}|^{1/2}} \exp\left\{ -\frac{1}{2} (\mathbf{x}^{(i)} - \boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}^{(i)} - \boldsymbol{\mu}) \right\}. \end{align*}\]
We maximize this over \(\boldsymbol{\mu}\) and \(\boldsymbol{\Sigma}\). It is easier to maximize a sum rather than a product, so we take the log-likelihood \(\ell(\boldsymbol{\mu}, \boldsymbol{\Sigma}) = \log \mathcal{L}(\boldsymbol{\mu}, \boldsymbol{\Sigma})\):
\[\begin{align*} \ell(\boldsymbol{\mu}, \boldsymbol{\Sigma}) &= \sum_{i=1}^{N} \left[ -\frac{D}{2}\log(2\pi) - \frac{1}{2}\log|\boldsymbol{\Sigma}| - \frac{1}{2} (\mathbf{x}^{(i)} - \boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}^{(i)} - \boldsymbol{\mu}) \right] \\ &= -\frac{N}{2}\log(2\pi) - \frac{N}{2}\log|\boldsymbol{\Sigma}| - \frac{1}{2}\sum_{i=1}^{N} (\mathbf{x}^{(i)} - \boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x}^{(i)} - \boldsymbol{\mu}). \end{align*}\]
To maximize \(\ell\) with respect to \(\boldsymbol{\mu}\), we set its derivative with respect to \(\boldsymbol{\mu}\) to zero. It is possible to show, either with vector calculus laws or by writing out the matrix-vector multiplications, that \[ \nabla_{\boldsymbol{\mu}} \ell = \sum_{i=1}^{N} \boldsymbol{\Sigma}^{-1}(\mathbf{x}^{(i)} - \boldsymbol{\mu}) = 0. \] Solving gives \[ \hat{\boldsymbol{\mu}}_{\textrm{MLE}} = \frac{1}{N} \sum_{i=1}^{N} \mathbf{x}^{(i)}. \] So the maximum-likelihood estimate of the mean is the sample mean (or empirical mean) of the observed vectors.
To maximize \(\ell\) with respect to \(\boldsymbol{\Sigma}\), we again set the appropriate derivatives to zero. The derivation is omitted (and beyond the scope of this course), but it yields \[ \hat{\boldsymbol{\Sigma}}_{\textrm{MLE}} = \frac{1}{N} \sum_{i=1}^{N} (\mathbf{x}^{(i)} - \hat{\boldsymbol{\mu}})(\mathbf{x}^{(i)} - \hat{\boldsymbol{\mu}})^\top. \] This is the sample covariance (or empirical covariance) matrix, where the \((j,\ell)\) entry is the sample covariance between the \(j\)-th and \(\ell\)-th features.
In the next chapter we use this distribution as the class-conditional model in multivariate Gaussian Discriminant Analysis.
3.1 A Multivariate Gaussian for Maple Leaves
To make this concrete, we fit a bivariate Gaussian to the three maple leaves in the dataset.
| # | Leaf Width | Leaf Height | Leaf Type |
|---|---|---|---|
| 5 | 13.0 cm | 11.0 cm | |
| 6 | 11.0 cm | 9.0 cm | |
| 7 | 9.0 cm | 14.0 cm |
With \(N=3\) and \(D=2\), we have the following MLE estimate for the mean:
\[ \hat{\boldsymbol{\mu}} = \frac{1}{3}\left[\begin{bmatrix}13\\11\end{bmatrix}+\begin{bmatrix}11\\9\end{bmatrix}+\begin{bmatrix}9\\14\end{bmatrix}\right] \approx \begin{bmatrix}11.0\\11.3\end{bmatrix}\text{ cm.} \]
For the sample variance, we again use the MLE estimate formula from above:
\[\begin{align*} \hat{\boldsymbol{\Sigma}} &= \frac{1}{3}\left[ \left(\mathbf{x}^{(1)}-\hat{\boldsymbol{\mu}}\right)\left(\mathbf{x}^{(1)}-\hat{\boldsymbol{\mu}}\right)^\top + \left(\mathbf{x}^{(2)}-\hat{\boldsymbol{\mu}}\right)\left(\mathbf{x}^{(2)}-\hat{\boldsymbol{\mu}}\right)^\top + \left(\mathbf{x}^{(3)}-\hat{\boldsymbol{\mu}}\right)\left(\mathbf{x}^{(3)}-\hat{\boldsymbol{\mu}}\right)^\top \right] \\ &= \frac{1}{3}\left[ \begin{bmatrix}4 & -\tfrac{2}{3}\\-\tfrac{2}{3} & \tfrac{1}{9}\end{bmatrix} + \begin{bmatrix}0 & 0\\0 & \tfrac{49}{9}\end{bmatrix} + \begin{bmatrix}4 & -\tfrac{16}{3}\\-\tfrac{16}{3} & \tfrac{64}{9}\end{bmatrix} \right] \\ &\approx \begin{bmatrix}2.67 & -2.00\\-2.00 & 4.22\end{bmatrix}. \end{align*}\] The negative off-diagonal reflects a mild negative correlation between width and height in this small sample.
The figure below shows the three maple data points (orange) with the level sets of the fitted Gaussian overlaid. The red dot marks \(\hat{\boldsymbol{\mu}}\).
Fitted Gaussian for maple leaves
Question: Apply the same procedure to the four oak leaves with (width, height) = (7, 12), (9, 6), (10, 18), (5, 10) cm. Compute \(\hat{\boldsymbol{\mu}}\) and \(\hat{\boldsymbol{\Sigma}}\).
Answer:
Sample mean: \[ \hat{\boldsymbol{\mu}} = \frac{1}{4}\left[\begin{bmatrix}7\\12\end{bmatrix}+\begin{bmatrix}9\\6\end{bmatrix}+\begin{bmatrix}10\\18\end{bmatrix}+\begin{bmatrix}5\\10\end{bmatrix}\right] = \frac{1}{4}\begin{bmatrix}31\\46\end{bmatrix} = \begin{bmatrix}7.75\\11.5\end{bmatrix}\text{ cm.} \]
Sample covariance: \[ \hat{\boldsymbol{\Sigma}} = \frac{1}{4}\sum_{i=1}^{4}(\mathbf{x}^{(i)}-\hat{\boldsymbol{\mu}})(\mathbf{x}^{(i)}-\hat{\boldsymbol{\mu}})^\top \approx \begin{bmatrix}3.69 & 2.88\\2.88 & 18.75\end{bmatrix}. \] The positive off-diagonal indicates that wider oak leaves also tend to be taller.
4 Understanding the Covariance Matrix
The figures earlier show that the contours of a bivariate Gaussian are ellipses—tilted when \(\sigma_{12} \neq 0\), axis-aligned when \(\boldsymbol{\Sigma}\) is diagonal. But how do we know the contours look like that, and what determines their orientation and shape? We answer this by focusing on the quadratic form that appears in the exponent of the Gaussian.
Definition: A quadratic form is a scalar-valued expression of the form \(\mathbf{x}^\top A \mathbf{x}\), where \(A\) is a symmetric matrix. The exponent of the Gaussian, \((\mathbf{x} - \boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x} - \boldsymbol{\mu})\), is a quadratic form in the centered variable \((\mathbf{x} - \boldsymbol{\mu})\) with \(A = \boldsymbol{\Sigma}^{-1}\).
Recall that contour lines (or level sets) of a function are the regions where the function takes a constant value. In our case the function is the density \(p(\mathbf{x})\). So a contour of equal density is the set of all \(\mathbf{x}\) for which \(p(\mathbf{x}) = C\) for some constant \(C\). The Gaussian density is \[p(\mathbf{x}) \propto \exp\bigl(-\frac{1}{2}(\mathbf{x} - \boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x} - \boldsymbol{\mu})\bigr),\] so \(p(\mathbf{x})\) is constant if and only if the exponent \((\mathbf{x} - \boldsymbol{\mu})^\top \boldsymbol{\Sigma}^{-1} (\mathbf{x} - \boldsymbol{\mu})\) is constant. In other words, contours of equal density are exactly the level sets of this quadratic form.
Again, we will build intuition in \(D=2\), but reasoning mathematically helps us understand the role of \(\boldsymbol{\Sigma}\) more generally. In 2D, the contours of the quadratic form \(\mathbf{x}^\top \boldsymbol{\Sigma}^{-1} \mathbf{x}\) are ellipses; in higher dimensions they are ellipsoids.
To further simplify, take the mean to be the zero vector, \(\boldsymbol{\mu} = \mathbf{0}\) (we can do this because \(\boldsymbol{\mu}\) just shifts the location of the contours without changing its shape, as shown in Figure 1). Then the exponent reduces to \(\mathbf{x}^\top \boldsymbol{\Sigma}^{-1} \mathbf{x}\), and the contours are the sets \(\bigl\{ \mathbf{x} : \mathbf{x}^\top \boldsymbol{\Sigma}^{-1} \mathbf{x} = C \bigr\}\) for constants \(C \geq 0\). So understanding the shape and orientation of the contours amounts to understanding the behaviour of the quadratic form \(\mathbf{x}^\top \boldsymbol{\Sigma}^{-1} \mathbf{x}\).
Definition: The principal axes of an ellipse are its axes of symmetry. These are the directions along which the ellipse is widest and narrowest. Each such direction is a principal direction. In 2D an ellipse has two principal axes, which are always perpendicular to each other.
It turns out that the eigenvectors of \(\boldsymbol{\Sigma}\) (and \(\boldsymbol{\Sigma}^{-1}\), as they share the same eigenvectors) are precisely the directions of the principal axes of the contour ellipsoid. The semi-axes of the ellipse lie along these directions. The eigenvalues of \(\boldsymbol{\Sigma}\) tell us how spread out the Gaussian is along each of those axes: larger eigenvalue \(\lambda_j\) means more variance (and a longer semi-axis) in the direction of the corresponding eigenvector \(\mathbf{q}_j\); smaller eigenvalue means a shorter semi-axis. So the eigenvectors answer “which directions are the axes of the ellipse?” and the eigenvalues answer “how long is each semi-axis?” (semi-axis length in direction \(\mathbf{q}_j\) is proportional to \(\sqrt{\lambda_j}\)).
Definition: An eigenvector of a square matrix \(A\) is a nonzero vector \(\mathbf{q}\) satisfying \[A\mathbf{q} = \lambda \mathbf{q}\] for some scalar \(\lambda\), called the corresponding eigenvalue.
Question: Is \(\mathbf{v} = \begin{bmatrix} 1 \\ 1 \end{bmatrix}\) an eigenvector of \(\boldsymbol{A} = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}\)? If so, what is the corresponding eigenvalue?
Answer:
\[ \boldsymbol{A}\mathbf{v} = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}\begin{bmatrix} 1 \\ 1 \end{bmatrix} = \begin{bmatrix} 3 \\ 3 \end{bmatrix} = 3\begin{bmatrix} 1 \\ 1 \end{bmatrix} = 3\mathbf{v}. \] Since \(\boldsymbol{A}\mathbf{v} = 3\mathbf{v}\), yes: \(\mathbf{v}\) is an eigenvector of \(\boldsymbol{A}\) with eigenvalue \(\lambda = 3\).
Question: Show that \(\boldsymbol{\Sigma}\) and \(\boldsymbol{\Sigma}^{-1}\) share the same eigenvectors. What is the eigenvalue of \(\boldsymbol{\Sigma}^{-1}\) corresponding to an eigenvector with eigenvalue \(\lambda\) in \(\boldsymbol{\Sigma}\)?
Answer:
Suppose \(\mathbf{q}\) is an eigenvector of \(\boldsymbol{\Sigma}\) with eigenvalue \(\lambda > 0\), so \(\boldsymbol{\Sigma}\mathbf{q} = \lambda\mathbf{q}\). Multiply both sides on the left by \(\boldsymbol{\Sigma}^{-1}\): \[ \mathbf{q} = \lambda\,\boldsymbol{\Sigma}^{-1}\mathbf{q}. \] Dividing by \(\lambda\): \[ \boldsymbol{\Sigma}^{-1}\mathbf{q} = \frac{1}{\lambda}\mathbf{q}. \] So \(\mathbf{q}\) is an eigenvector of \(\boldsymbol{\Sigma}^{-1}\) with eigenvalue \(1/\lambda\). Since this holds for every eigenvector of \(\boldsymbol{\Sigma}\), the two matrices share the same eigenvectors; only their eigenvalues differ (they are reciprocals of each other).
We show the relationship between eigenvectors and contour shapes first for the simpler case of a diagonal \(\boldsymbol{\Sigma}\), where the geometry is transparent, then extend to the general case using the spectral decomposition of \(\boldsymbol{\Sigma}\).
4.1 Diagonal \(\boldsymbol{\Sigma}\): Axis-Aligned Contours
If \(\boldsymbol{\Sigma}\) is diagonal, then as shown in Figure 3, for \(\boldsymbol{\Sigma} = \mathrm{diag}(\sigma_1^2, \sigma_2^2)\) the ellipse axes always align with \(x_1\) and \(x_2\), regardless of the values of \(\sigma_1^2\) and \(\sigma_2^2\).
In other words, the principal directions of the ellipse are the standard basis vectors \(\begin{bmatrix} 1 &0 \end{bmatrix}^\top\) and \(\begin{bmatrix} 0 & 1 \end{bmatrix}^\top\). These are precisely the eigenvectors of \(\boldsymbol{\Sigma}\). The corresponding eigenvalues, \(\sigma_1^2\) and \(\sigma_2^2\), give the variance along each axis: larger \(\sigma_j^2\) means the distribution is more spread out in the \(x_j\) direction, producing a longer semi-axis of length proportional to \(\sigma_j\).
Question: The contour plot below shows level sets of a bivariate Gaussian with \(\boldsymbol{\mu} = \mathbf{0}\). Which of the following could be its covariance matrix \(\boldsymbol{\Sigma}\)?
- \(\begin{bmatrix} 2 & 0 \\ 0 & 1 \end{bmatrix}\) \(\quad\) (B) \(\begin{bmatrix} 1 & 0 \\ 0 & 2 \end{bmatrix}\) \(\quad\) (C) \(\begin{bmatrix} 1 & 0.3 \\ 0.3 & 2 \end{bmatrix}\)
Answer:
(A) is the only valid choice.
Option (A) is diagonal with \(\sigma_1^2 = 2 > \sigma_2^2 = 1\), so the contours are axis-aligned and wider along \(x_1\) (semi-axis proportional to \(\sqrt{2}\)) than along \(x_2\) (semi-axis proportional to \(1\)), which is consistent with the plot.
Option (B) is diagonal with \(\sigma_1^2 = 1 < \sigma_2^2 = 2\), so it would produce axis-aligned contours wider along \(x_2\), not \(x_1\).
Option (C) has a nonzero off-diagonal, which would tilt the contours away from the axes. It is inconsistent with the axis-aligned ellipses shown.
4.2 Non-Diagonal \(\boldsymbol{\Sigma}\): Change of Basis
When \(\boldsymbol{\Sigma}\) is not diagonal, the contours are tilted ellipses (in 2D) or ellipsoids (in higher dimensions). The idea of the “tilt” is key. As shown in Figure 7 (left), the tilted ellipse becomes axis-aligned in a rotated coordinate system with axes \(\tilde{x}_1\) and \(\tilde{x}_2\). A change of basis converts the tilted ellipse into the axis-aligned one.
Definition: A change of basis expresses the same vector in a different coordinate system. If \(\mathbf{Q}\) is an orthogonal matrix whose columns form a new basis, then \(\tilde{\mathbf{x}} = \mathbf{Q}^\top \mathbf{x}\) gives the coordinates of \(\mathbf{x}\) in that basis. Because \(\mathbf{Q}\) is orthogonal (\(\mathbf{Q}^\top \mathbf{Q} = \mathbf{I}\)), this is a rotation (lengths and angles are preserved), and we can recover \(\mathbf{x} = \mathbf{Q}\tilde{\mathbf{x}}\).
Gaussian contours in an eigenvector basis
In our case, the specifics of the change of basis is actually given by the spectral theorem
Definition: The Spectral Theorem tells us that any real symmetric matrix \(\boldsymbol{\Sigma}\) can be decomposed as \[\boldsymbol{\Sigma} = \mathbf{Q}\boldsymbol{\Lambda}\mathbf{Q}^\top,\] where \(\mathbf{Q}\) is an orthogonal matrix whose columns \(\mathbf{q}_1,\ldots,\mathbf{q}_D\) are the eigenvectors of \(\boldsymbol{\Sigma}\), and \(\boldsymbol{\Lambda} = \mathrm{diag}(\lambda_1,\ldots,\lambda_D)\) contains the corresponding eigenvalues on its diagonal.
The decomposition \(\boldsymbol{\Sigma} = \mathbf{Q}\boldsymbol{\Lambda}\mathbf{Q}^\top\) has a geometric interpretation as a change of basis. Define rotated coordinates \(\tilde{\mathbf{x}} = \mathbf{Q}^\top \mathbf{x}\), which align the axes with the eigenvectors of \(\boldsymbol{\Sigma}\). In these coordinates the quadratic form simplifies to \[ \mathbf{x}^\top \boldsymbol{\Sigma}^{-1} \mathbf{x} = \tilde{\mathbf{x}}^\top \boldsymbol{\Lambda}^{-1} \tilde{\mathbf{x}} = \sum_{j=1}^{D} \frac{\tilde{x}_j^2}{\lambda_j}, \] which is an axis-aligned ellipsoid in \(\tilde{\mathbf{x}}\) with semi-axes proportional to \(\sqrt{\lambda_j}\). Rotating back to the original coordinates, the ellipsoid’s axes point along the eigenvectors \(\mathbf{q}_j\), with lengths proportional to \(\sqrt{\lambda_j}\). The figure above illustrates this: the same bivariate Gaussian is shown in the original coordinates (tilted contours) and in coordinates aligned with the eigenvectors (axis-aligned contours).
In summary, the off-diagonal entries of \(\boldsymbol{\Sigma}\) encode correlation between features, which appears geometrically as a tilt of the contour ellipsoid. The eigenvectors of \(\boldsymbol{\Sigma}\) determine the orientation of that tilt, and the eigenvalues determine the spread along each principal direction. Together they fully characterize the shape of the Gaussian’s contours.
5 Summary
The multivariate Gaussian distribution is an extension of the univariate Gaussian (Normal) distribution. Just like before, we can learn this distribution using Maximum Likelihood estimation. The parameters \(\boldsymbol{\mu}\) and \(\boldsymbol{\Sigma}\) generalize the univariate mean and variance, where \(\boldsymbol{\mu}\) places the distribution in space, and \(\boldsymbol{\Sigma}\) encodes the distribution’s spread and “shape”. The spectral decomposition of \(\boldsymbol{\Sigma}\) tells us more about the distribution: the eigenvectors of \(\boldsymbol{\Sigma}\) give its orientation, and the eigenvalues give its spread along each principal direction.
This chapter illustrates two of the fundamental ideas introduced at the start of the book. The covariance matrix and its contour ellipsoids are a direct instance of Fundamental Idea #4: ML Describes Geometric Processes. The structure of a distribution is encoded in the geometry of its level sets, and the spectral decomposition reveals that geometry precisely. At the same time, GDA is a model built entirely on Fundamental Idea #5: ML Demands a Probabilistic Lens. Instead of drawing a boundary directly, we model the data-generating process as a mixture of Gaussians and let Bayes’ rule produce the classifier. In fact, the same mathematics, including eigenvectors, eigenvalues, and the spectral decomposition, will reappear in Principal Component Analysis (PCA) unit.