Probabilistic Conditional Models
1 Introduction
On the previous page, we discussed how to learn a probabilistic model that approximates the data-generating distribution. We showed how MLE and MAP estimate the parameters of that distribution from data. Throughout, we focused on one question, “What is the probability of winning our next game?”, which we modelled with one Bernoulli parameter.
This page extends that example by introducing a second variable, which holds information about the game that is available before it is played. This allows us to ask a more informative question. What is the probability of winning, given what we know about the game? We refer to a model that answers questions of this form as a probabilistic conditional model. Whereas the model on the previous page described the distribution of the outcome alone, a probabilistic conditional model describes how the distribution of the outcome changes with the information we observe.
We develop these ideas in three stages. First, we introduce a binary feature indicating whether each game was played at home or away. We model the joint distribution of the feature and the outcome, estimate its parameters using maximum likelihood, and apply Bayes’ rule to compute the probability of winning a home game. Second, we introduce a continuous feature, the opponent’s skill rating. The approach remains the same, but Gaussian distributions replace the Bernoulli distributions for the feature given the outcome, which allows the model to make predictions for ratings not observed in the data. Finally, we consider multiple features simultaneously and show that the number of parameters grows exponentially with the number of features. The next page addresses this limitation.
2 A Binary Feature: Home vs. Away
Suppose we also record whether each game was played at home or away. This additional information allows us to ask a richer question:
What is the probability that our team wins a game, given that the game was played at home?
Let’s say that we took the same 20 games from the previous page (14 wins, 6 losses overall), and split them by whether the game was played at home:
| Win (\(Y = 1\)) | Loss (\(Y = 0\)) | |
|---|---|---|
| Home (\(X = 1\)) | 8 | 2 |
| Away (\(X = 0\)) | 6 | 4 |
What is the probability that our team wins a home game?
Intuitively, our team won \(8\) of the \(10\) home games, so the answer is \[ \begin{aligned} & \frac{8}{8 + 2} = \frac{8}{10} \end{aligned} \] As before, we want to justify this rather than assert it.
2.1 The Model
Modelling this new scenario requires two binary random variables. We will use \(Y \in \{0, 1\}\) to model the outcome of a game. \(Y = 1\) if our team won and \(Y = 0\) otherwise. We will use \(X \in \{0, 1\}\) to model where the game was played. \(X = 1\) for the home field and \(X = 0\) for the away field.
To capture how \(X\) and \(Y\) relate, we will model their joint distribution \(P(X, Y)\). Since both variables are binary, the pair \((X, Y)\) has four possible outcomes. Their probabilities must sum to \(1\), so three parameters are enough to fully define the joint distribution.
Rather than specifying the four probabilities directly, it is convenient to factor the joint distribution using the product rule.
Definition: For random variables \(A\) and \(B\), the product rule states that \[ \begin{aligned} & P(A, B) = P(A) \, P(B \mid A) \end{aligned} \]
Applying the product rule to \(X\) and \(Y\) gives \[ \begin{aligned} & P(X, Y) = P(Y) \, P(X \mid Y) \end{aligned} \] We can then model each factor using Bernoulli distributions.
First, we model \(P(Y)\) as a Bernoulli distribution: \[ \begin{aligned} & Y \sim \mathrm{Bernoulli}(\pi) \end{aligned} \] where the parameter \(\pi\) represents the probability that our team wins a game.
Second, we model \(P(X \mid Y)\). Since \(Y\) has two possible values, we use a separate Bernoulli distribution for each value of \(Y\): \[ \begin{aligned} & X \mid Y = 0 \sim \mathrm{Bernoulli}(\theta_0), \qquad X \mid Y = 1 \sim \mathrm{Bernoulli}(\theta_1) \end{aligned} \] where the parameter \(\theta_1\) represents the probability that a game was played at home given that our team won, and \(\theta_0\) represents the probability that a game was played at home given that our team lost.
Together, these three parameters fully define the joint distribution. Let’s use \(\theta\) to denote the full set of model parameters: \[ \begin{aligned} & \theta = \{\pi, \theta_0, \theta_1\} \end{aligned} \] Writing out each probability explicitly, our model is: \[ \begin{aligned} & P(Y = 1 \,;\, \theta) = \pi, \\ & P(X = 1 \mid Y = 0 \,;\, \theta) = \theta_0, \\ & P(X = 1 \mid Y = 1 \,;\, \theta) = \theta_1 \end{aligned} \]
Using the product rule, we can compute each entry of the joint distribution from the model parameters: \[ \begin{aligned} & P(X = x, Y = y \,;\, \theta) = P(Y = y \,;\, \theta) \, P(X = x \mid Y = y \,;\, \theta) \end{aligned} \] Table 2 shows the result for all four outcomes.
| Win (\(Y = 1\)) | Loss (\(Y = 0\)) | |
|---|---|---|
| Home (\(X = 1\)) | \(\pi \, \theta_1\) | \((1 - \pi) \, \theta_0\) |
| Away (\(X = 0\)) | \(\pi \, (1 - \theta_1)\) | \((1 - \pi) \, (1 - \theta_0)\) |
Some of the exercises below also use the law of total probability.
Definition: For random variables \(A\) and \(B\), where \(B\) takes values in a finite set \(\mathcal{B}\), the law of total probability states that \[ \begin{aligned} & P(A) = \sum_{b \in \mathcal{B}} P(A, B = b) \end{aligned} \]
Answer:
\[ \begin{aligned} & P(X = 0 \mid Y = 1) = 1 - \theta_{1} \end{aligned} \]
Answer:
First, we can introduce \(Y\) into the expression by using the law of total probability, summing over \(Y \in \{0,1\}\):
\[ \begin{aligned} & P(X = 1) = \sum_{y \in \{ 0, 1 \}} P(X = 1, Y = y) \\ & = P(X = 1, Y = 0) + P(X = 1, Y = 1) \end{aligned} \]
Next, we decompose each joint probability of \(X\) and \(Y\) using the product rule:
\[ \begin{aligned} & P(X = 1) = P(Y = 0) \, P(X = 1 \mid Y = 0) \\ & \qquad\qquad\qquad + P(Y = 1) \, P(X = 1 \mid Y = 1) \end{aligned} \]
Finally, we substitute in the model parameters:
\[ \begin{aligned} & P(X = 1) = (1 - \pi) \, \theta_0 + \pi \, \theta_1 \end{aligned} \]
Answer:
By the definition of conditional probability:
\[ \begin{aligned} & P(Y = 1 \mid X = 0) = \frac{P(X = 0, Y = 1)}{P(X = 0)} \end{aligned} \]
The numerator is the probability that the game was played on the away field and our team won the game. Reading from Table 2, it is \[ \begin{aligned} & P(X = 0, Y = 1) = \pi \, (1 - \theta_1) \end{aligned} \]
The denominator is the probability that the game was played on the away field, regardless of the outcome of the game. We can calculate the denominator using the law of total probability as follows: \[ \begin{aligned} & P(X = 0) = P(X = 0, Y = 0) + P(X = 0, Y = 1) \\ & = (1 - \pi) \, (1 - \theta_0) + \pi \, (1 - \theta_1) \end{aligned} \]
Substituting the numerator and denominator gives
\[ \begin{aligned} & P(Y = 1 \mid X = 0) = \frac{\pi \, (1 - \theta_1)}{(1 - \pi) \, (1 - \theta_0) + \pi \, (1 - \theta_1)} \end{aligned} \]
2.2 Computing the MLE of the Parameters
Before deriving the MLE formally, let’s compute it intuitively. As on the previous page, the MLE of each parameter is the empirical frequency of the corresponding outcome in our data. Try computing the MLE of \(\pi\) using our data in the exercise below.
Answer:
The parameter \(\pi\) is the probability that our team wins a game. Our team won \(8\) home games and \(6\) away games, which is \(14\) of the \(20\) games, so \[ \begin{aligned} & \hat{\pi}_{\text{MLE}} = \frac{14}{20} = \frac{7}{10} \end{aligned} \]
Next, let’s compute the MLE of \(\theta_0\). The parameter \(\theta_0\) is the probability that a game was played at home given that our team lost. So we only look at the games our team lost. Our team lost \(2\) home games and \(4\) away games, so \(2\) of its \(6\) losses were at home: \[ \begin{aligned} & \hat{\theta}_{0, \text{MLE}} = \frac{2}{6} = \frac{1}{3} \end{aligned} \]
Answer:
The parameter \(\theta_1\) is the probability that a game was played at home given that our team won. So we only look at the games our team won. Our team won \(8\) home games and \(6\) away games, so \(8\) of its \(14\) wins were at home: \[ \begin{aligned} & \hat{\theta}_{1, \text{MLE}} = \frac{8}{14} = \frac{4}{7} \end{aligned} \]
In the next section, we will derive these estimates formally and show that they are indeed the MLE.
2.3 Deriving the MLE of the Parameters
Let’s derive the MLE of the model parameters. Now our data set keeps track of the values of two random variables for each data point. \[ \begin{aligned} & \mathcal{D} = \left\{ \left( x^{(1)}, y^{(1)} \right), \left( x^{(2)}, y^{(2)} \right), \left( x^{(3)}, y^{(3)} \right), \ldots, \left( x^{(N)}, y^{(N)} \right) \right\} \end{aligned} \]
We can follow the same procedure for deriving the MLE from the previous page:
- Derive the likelihood of the data given the parameter(s).
- Derive the log-likelihood of the data given the parameter(s).
- Take the derivative of the log-likelihood w.r.t. the parameter(s).
- Set the derivative to \(0\) and solve for the parameter(s).
Step 0: Derive the Likelihood of One Data Point
First, we need to write down an expression for \(P\left(x^{(i)}, y^{(i)} \,;\, \theta \right)\), the probability of the \(i\)-th data point. As shown above, we can factor this joint probability using the product rule: \[ \begin{aligned} & P\left( x^{(i)}, y^{(i)} \,;\, \theta \right) = P\left( y^{(i)} \,;\, \theta \right) \, P\left( x^{(i)} \mid y^{(i)} \,;\, \theta \right) \end{aligned} \] Now, we will derive an expression for each factor. Recall from the previous page that we can write a probability compactly using the exponent trick. Let’s start with a compact expression for \(P\left( y^{(i)} \,;\, \theta \right)\). When our team wins the game, the probability is \(\pi\): \[ \begin{aligned} & P\left( y^{(i)} = 1 \,;\, \theta \right) = \pi \end{aligned} \]
When our team loses the game, the probability is \(1 - \pi\): \[ \begin{aligned} & P\left( y^{(i)} = 0 \,;\, \theta \right) = 1 - \pi \end{aligned} \]
Combining the two cases using the exponent trick, we get: \[ \begin{aligned} & P\left( y^{(i)} \,;\, \theta \right) = \pi^{y^{(i)}} (1 - \pi)^{(1 - y^{(i)})} \end{aligned} \]
When \(y^{(i)} = 1\), the second factor has exponent \(0\), so it equals \(1\) and leaves \(\pi\). When \(y^{(i)} = 0\), the first factor has exponent \(0\), so it equals \(1\) and leaves \(1 - \pi\).
We can express \(P\left( x^{(i)} \mid y^{(i)} \,;\, \theta \right)\) compactly in a similar way. Try deriving it yourself in the exercise below.
Answer:
When our team loses the game (\(y^{(i)} = 0\)), \(X\) follows a Bernoulli distribution with parameter \(\theta_0\): \[ \begin{aligned} & P\left( x^{(i)} = 1 \mid y^{(i)} = 0 \,;\, \theta \right) = \theta_0 \\ & P\left( x^{(i)} = 0 \mid y^{(i)} = 0 \,;\, \theta \right) = 1 - \theta_0 \end{aligned} \] We combine the two cases using the exponent trick: \[ \begin{aligned} & P\left( x^{(i)} \mid y^{(i)} = 0 \,;\, \theta \right) = (\theta_{0})^{x^{(i)}} (1 - \theta_{0})^{(1 - x^{(i)})} \end{aligned} \] Similarly, the compact expression for when our team wins the game (\(y^{(i)} = 1\)) is as follows: \[ \begin{aligned} & P\left( x^{(i)} \mid y^{(i)} = 1 \,;\, \theta \right) = (\theta_{1})^{x^{(i)}} (1 - \theta_{1})^{(1 - x^{(i)})} \end{aligned} \] Next, we combine the two compact expressions using the exponent trick: \[ \begin{aligned} & P\left( x^{(i)} \mid y^{(i)} \,;\, \theta \right) \\ & = \left[ (\theta_{0})^{x^{(i)}} (1 - \theta_{0})^{(1 - x^{(i)})} \right]^{(1 - y^{(i)})} \left[ (\theta_{1})^{x^{(i)}} (1 - \theta_{1})^{(1 - x^{(i)})} \right]^{y^{(i)}} \end{aligned} \]
Therefore, the probability of the \(i\)-th data point \(P\left( x^{(i)}, y^{(i)} \,;\, \theta \right)\) can be expressed as follows:
\[ \begin{aligned} & P\left( x^{(i)}, y^{(i)} \,;\, \theta \right) \\ & = \underbrace{\left[ \pi^{y^{(i)}} (1 - \pi)^{(1 - y^{(i)})} \right]}_{P\left( y^{(i)} \right)} \\ & \quad \cdot \underbrace{ \left[ (\theta_{0})^{x^{(i)}} (1 - \theta_{0})^{(1 - x^{(i)})} \right]^{(1 - y^{(i)})} \left[ (\theta_{1})^{x^{(i)}} (1 - \theta_{1})^{(1 - x^{(i)})} \right]^{y^{(i)}} }_{P\left( x^{(i)} \mid y^{(i)} \right)} \end{aligned} \]
Step 1: Derive the Likelihood of the Data Set
Since the games are i.i.d. (defined in Understanding the Likelihood Function on the previous page), the likelihood is a product over the games: \[ \begin{aligned} & L(\theta) = P\left( \mathcal{D} \,;\, \theta \right) = \prod_{i=1}^{N} P\left(x^{(i)}, y^{(i)} \,;\, \theta \right) \end{aligned} \]
Plugging our expression for the probability of the \(i\)-th data point into this product, we get: \[ \begin{aligned} L(\theta) &= P\left( \mathcal{D} \,;\, \theta \right) \\[0.5em] &= \prod_{i=1}^{N} P\left( x^{(i)}, y^{(i)} \,;\, \theta \right) \\[0.5em] &= \prod_{i=1}^{N} P\left( y^{(i)} \,;\, \theta \right) \, P\left( x^{(i)} \mid y^{(i)} \,;\, \theta \right) \\[0.5em] &= \prod_{i=1}^{N} \left[\pi^{y^{(i)}} (1 - \pi)^{(1 - y^{(i)})} \right] \left[ (\theta_{0})^{x^{(i)}} (1 - \theta_{0})^{(1 - x^{(i)})} \right]^{(1 - y^{(i)})} \left[ (\theta_{1})^{x^{(i)}} (1 - \theta_{1})^{(1 - x^{(i)})} \right]^{y^{(i)}} \end{aligned} \]
Step 2: Derive the Log-Likelihood
Next, we calculate the log-likelihood \(\ell(\theta)\). Try deriving the expression yourself in the exercise below.
Question: Starting from the likelihood function \(L(\theta)\) above, show that the log-likelihood can be written as follows: \[ \begin{aligned} \ell (\theta) &= \sum_{i = 1}^{N} \left[ y^{(i)} \log(\pi) + \left( 1 - y^{(i)} \right) \log(1 - \pi) \right] \\[0.5em] & \quad + \sum_{i = 1}^{N} \left[ \left( 1 - y^{(i)} \right) \left( x^{(i)} \log(\theta_{0}) + \left( 1 - x^{(i)} \right) \log(1 - \theta_{0}) \right) \right] \\[0.5em] & \quad + \sum_{i = 1}^{N} \left[ y^{(i)} \left( x^{(i)} \log(\theta_{1}) + \left( 1 - x^{(i)} \right) \log(1 - \theta_{1}) \right) \right] \end{aligned} \]
Answer:
First, the log turns the product over data points into a sum: \[ \begin{aligned} \ell(\theta) &= \log L(\theta) \\[0.5em] &= \sum_{i=1}^{N} \log \left( \left[\pi^{y^{(i)}} (1 - \pi)^{(1 - y^{(i)})} \right] \left[ (\theta_{0})^{x^{(i)}} (1 - \theta_{0})^{(1 - x^{(i)})} \right]^{(1 - y^{(i)})} \left[ (\theta_{1})^{x^{(i)}} (1 - \theta_{1})^{(1 - x^{(i)})} \right]^{y^{(i)}} \right) \\[0.5em] &= \sum_{i=1}^{N} \Big[ y^{(i)} \log(\pi) + \left( 1 - y^{(i)} \right) \log(1 - \pi) \\[0.5em] & \qquad\qquad + \left( 1 - y^{(i)} \right) \left( x^{(i)} \log(\theta_{0}) + \left( 1 - x^{(i)} \right) \log(1 - \theta_{0}) \right) \\[0.5em] & \qquad\qquad + y^{(i)} \left( x^{(i)} \log(\theta_{1}) + \left( 1 - x^{(i)} \right) \log(1 - \theta_{1}) \right) \Big] \end{aligned} \]
Steps 3 and 4: Take the Derivative and Solve
The log-likelihood nicely decomposes into three terms, with each model parameter appearing in exactly one term. This means we can optimize each parameter independently of the others.
To find the MLE of each parameter, we take the derivative of its term only, set it to \(0\), and solve. Try these steps yourself in the exercises below.
Answer:
The term containing \(\pi\) is \[ \begin{aligned} & \sum_{i = 1}^{N} \left[ y^{(i)} \log(\pi) + \left( 1 - y^{(i)} \right) \log(1 - \pi) \right] \\ & = N_{\text{win}} \log(\pi) + N_{\text{loss}} \log(1 - \pi) \end{aligned} \] where \(N_{\text{loss}}\) is the number of games our team lost. This has the same form as the log-likelihood on the previous page.
Taking the derivative with respect to \(\pi\) and setting it to \(0\): \[ \begin{aligned} & \frac{N_{\text{win}}}{\pi} - \frac{N_{\text{loss}}}{1 - \pi} = 0 \end{aligned} \]
Solving for \(\pi\): \[ \begin{aligned} & \hat{\pi}_{\text{MLE}} = \frac{N_{\text{win}}}{N_{\text{win}} + N_{\text{loss}}} \\ & = \frac{N_{\text{win}}}{N} \end{aligned} \]
Answer:
The term containing \(\theta_0\) is \[ \begin{aligned} & \sum_{i = 1}^{N} \left( 1 - y^{(i)} \right) \left( x^{(i)} \log(\theta_{0}) + \left( 1 - x^{(i)} \right) \log(1 - \theta_{0}) \right) \\ & = N_{\text{home, loss}} \log(\theta_0) + N_{\text{away, loss}} \log(1 - \theta_0) \end{aligned} \] where \(N_{\text{away, loss}}\) is the number of away games our team lost. The factor \(1 - y^{(i)}\) keeps only the games our team lost.
Taking the derivative with respect to \(\theta_0\) and setting it to \(0\): \[ \begin{aligned} & \frac{N_{\text{home, loss}}}{\theta_0} - \frac{N_{\text{away, loss}}}{1 - \theta_0} = 0 \end{aligned} \]
Solving for \(\theta_0\): \[ \begin{aligned} & \hat{\theta}_{0, \text{MLE}} = \frac{N_{\text{home, loss}}}{N_{\text{home, loss}} + N_{\text{away, loss}}} \\ & = \frac{N_{\text{home, loss}}}{N_{\text{loss}}} \end{aligned} \]
The derivation for \(\theta_1\) follows the same steps, with the factor \(y^{(i)}\) keeping only the games our team won.
Using our observed data, we can verify that these estimates indeed match the values we computed in the previous section: \[ \begin{aligned} & \hat{\pi}_{\text{MLE}} = \frac{7}{10} \quad\quad \hat{\theta}_{0, \text{MLE}} = \frac{1}{3} \quad\quad \hat{\theta}_{1, \text{MLE}} = \frac{4}{7} \end{aligned} \]
2.4 Inference
Having learned the model parameters, we are ready to answer our original question:
What is the probability that our team wins a home game?
We already know the answer intuitively. Our team won \(8\) of the \(10\) games played at home, so the probability is \(0.8\).
What if we want to compute this probability using the model parameters we estimated in the previous section? Doing this is less straightforward, since no model parameter directly corresponds to this probability. Instead, we will derive an expression for it using the model parameters only.
The probability we want is \(P\left( Y = 1 \mid X = 1 \right)\). Using Bayes’ rule, we can decompose it as follows: \[ \begin{aligned} & P\left( Y = 1 \mid X = 1 \right) = \frac{ P\left( Y = 1 \right) \, P\left( X = 1 \mid Y = 1 \right)}{P\left( X = 1 \right)} \end{aligned} \] The numerator is the product of two model parameters, \(\pi\) and \(\theta_1\). The denominator \(P\left( X = 1 \right)\), however, does not correspond to any model parameter, so let’s decompose it further.
First, we introduce \(Y\) using the law of total probability: \[ \begin{aligned} & P\left( X = 1 \right) = P\left( X = 1, Y = 0 \right) + P\left( X = 1, Y = 1 \right) \end{aligned} \] Next, we factor each term using the product rule: \[ \begin{aligned} & P\left( X = 1, Y = 0 \right) + P\left( X = 1, Y = 1 \right) \\ & = P\left( Y = 0 \right) \, P\left( X = 1 \mid Y = 0 \right) + P\left( Y = 1 \right) \, P\left( X = 1 \mid Y = 1 \right) \end{aligned} \] Substituting this back into Bayes’ rule gives an expression in which every factor corresponds to a model parameter: \[ \begin{aligned} &P\left( Y = 1 \mid X = 1 \right) \\ &= \frac{ P\left( Y = 1 \right) \, P\left( X = 1 \mid Y = 1 \right) } {P\left( Y = 0 \right) \, P\left( X = 1 \mid Y = 0 \right) + P\left( Y = 1 \right) \, P\left( X = 1 \mid Y = 1 \right)} \end{aligned} \] Writing each term using the model parameters, we get: \[ \begin{aligned} & P\left( Y = 1 \mid X = 1 \right) = \frac{ \pi \, \theta_1 }{ (1 - \pi) \, \theta_0 + \pi \, \theta_1 } \end{aligned} \]
Plugging in our estimates of the model parameters, we get: \[ \begin{aligned} & P\left( Y = 1 \mid X = 1 \right) = \frac{ \dfrac{7}{10} \cdot \dfrac{4}{7} }{ \dfrac{3}{10} \cdot \dfrac{1}{3} + \dfrac{7}{10} \cdot \dfrac{4}{7} } = 0.8 \end{aligned} \] This matches the answer we found intuitively.
Although this derivation seems lengthy for a probability we could compute by counting, it illustrates two important ideas. First, to capture how \(Y\) depends on \(X\), we modelled the joint distribution \(P(X, Y)\) by factoring it into \(P(Y)\) and \(P(X \mid Y)\), each of which we can estimate from data. Second, even when no model parameter directly answers our question, we can express the answer in terms of the model parameters using Bayes’ rule, the law of total probability, and the product rule. This approach applies more generally, including in settings where counting is not possible, as we will see in the next section.
3 A Continuous Feature: Opponent Rating
So far, \(X\) has been binary (home/away). Suppose we record a continuous feature instead, the opponent’s skill rating.
| Game | \(X\) (opponent rating) | \(Y\) (Win?) |
|---|---|---|
| 1 | 2.5 | 1 (win) |
| 2 | 1.5 | 1 (win) |
| 3 | 4.5 | 0 (loss) |
| 4 | 3.0 | 0 (loss) |
We are interested in answering the following question.
What is the probability that our team wins against an opponent whose rating is \(3.5\)?
This time we cannot derive an answer by counting. Worse, no data point has an opponent rating of \(3.5\). Fortunately, learning the joint distribution of \(X\) and \(Y\) does not require \(X\) to be binary. We keep the Bernoulli distribution for \(Y\) and only replace the Bernoulli distribution for \(X\) given \(Y\) with a continuous one. A natural choice is the Gaussian distribution.
3.1 The Model
The model is very similar to the one in the previous section. We model \(P(Y)\) as a Bernoulli distribution as before: \[ \begin{aligned} & Y \sim \mathrm{Bernoulli}(\pi) \end{aligned} \]
To model \(P(X \mid Y)\), let’s first recall the definition of a Gaussian distribution.
Definition: A continuous random variable \(X \in \mathbb{R}\) follows a Gaussian distribution with mean \(\mu \in \mathbb{R}\) and variance \(\sigma^2 > 0\) if its probability density function is \[ \begin{aligned} & \mathcal{N}(x \,;\, \mu, \sigma^2) = \frac{1}{\sqrt{2\pi\sigma^2}} \exp\left( -\frac{(x - \mu)^2}{2\sigma^2} \right) \end{aligned} \tag{1}\] We write this as \[ \begin{aligned} & X \sim \mathcal{N}\left( \mu, \sigma^2 \right) \end{aligned} \]
The mean \(\mu\) determines where the distribution is centred, and the variance \(\sigma^2\) determines how spread out it is.
For \(P(X \mid Y)\), we use a Gaussian distribution in place of each Bernoulli distribution: \[ \begin{aligned} & X \mid Y = 0 \sim \mathcal{N} \left(\mu_{0}, \sigma^2_{0} \right), \qquad X \mid Y = 1 \sim \mathcal{N} \left(\mu_{1}, \sigma^2_{1} \right) \end{aligned} \]
Together, these parameters fully define the joint distribution: \[ \begin{aligned} & \theta = \left\{ \pi, \mu_{0}, \sigma^2_{0}, \mu_{1}, \sigma^2_{1} \right\} \end{aligned} \] The parameter \(\pi\) is still the probability that our team wins a game. The parameters \(\mu_{0}\) and \(\sigma^2_{0}\) describe the distribution of opponent ratings given that our team lost, and \(\mu_{1}\) and \(\sigma^2_{1}\) describe the distribution of opponent ratings given that our team won.
3.2 Computing the MLE of the Parameters
As before, let’s compute the MLE intuitively before deriving it formally. The MLE of \(\pi\) is still the empirical frequency of wins in our data.
Start by computing the MLE of \(\pi\) in the exercise below.
Answer:
Recall that \(\pi\) is the probability that our team wins a game. Our team won \(2\) of the \(4\) games, so \[ \begin{aligned} & \hat{\pi}_{\text{MLE}} = \frac{2}{4} = 0.5 \end{aligned} \]
Now, let’s turn to \(\mu_0\) and \(\sigma_0^2\). These parameters describe the distribution of opponent ratings given that our team lost. We therefore only consider the games our team lost. Our team lost \(2\) games, against opponents rated \(4.5\) and \(3.0\). The MLE of \(\mu_0\) is the sample mean of these ratings: \[ \begin{aligned} & \hat{\mu}_{0, \text{MLE}} = \frac{4.5 + 3.0}{2} = 3.75 \end{aligned} \] The MLE of \(\sigma_0^2\) is the sample variance of these ratings: \[ \begin{aligned} \hat{\sigma}_{0, \text{MLE}}^2 & = \frac{(4.5 - 3.75)^2 + (3.0 - 3.75)^2}{2} \\[0.5em] & = \frac{0.5625 + 0.5625}{2} \\[0.5em] & = 0.5625 \end{aligned} \]
Answer:
The parameters \(\mu_1\) and \(\sigma_1^2\) describe the distribution of opponent ratings given that our team won. We therefore only consider the games our team won. Our team won \(2\) games, against opponents rated \(2.5\) and \(1.5\). The MLE of \(\mu_1\) is the sample mean of these ratings: \[ \begin{aligned} & \hat{\mu}_{1, \text{MLE}} = \frac{2.5 + 1.5}{2} = 2 \end{aligned} \] The MLE of \(\sigma_1^2\) is the sample variance of these ratings: \[ \begin{aligned} \hat{\sigma}_{1, \text{MLE}}^2 & = \frac{(2.5 - 2)^2 + (1.5 - 2)^2}{2} \\[0.5em] & = \frac{0.25 + 0.25}{2} \\[0.5em] & = 0.25 \end{aligned} \]
So opponents in games our team lost are rated \(3.75\) on average with variance \(0.5625\), while opponents in games our team won are rated \(2\) on average with variance \(0.25\).
Next, we will derive these estimates formally to confirm that they are the MLE.
3.3 Deriving the MLE of the Parameters
We follow the same procedure as in the previous section to derive the MLE of the parameters.
Step 0: Derive the Likelihood of One Data Point
As before, the probability of the outcome of the \(i\)-th game can be written compactly as follows: \[ \begin{aligned} & P\left( y^{(i)} \,;\, \theta \right) = \pi^{y^{(i)}} (1 - \pi)^{(1 - y^{(i)})} \end{aligned} \]
What is new is the second factor. Given the outcome of the game, the opponent rating now follows a Gaussian distribution rather than a Bernoulli distribution: \[ \begin{aligned} P\left( x^{(i)} \mid y^{(i)} \,;\, \theta \right) &= \mathcal{N}(x^{(i)} \,;\, \mu, \sigma^2) \\[0.5em] &= \frac{1}{\sqrt{2\pi\sigma^2}} \exp\left( -\frac{(x^{(i)} - \mu)^2}{2\sigma^2} \right) \end{aligned} \]
Step 1: Derive the Likelihood of the Data Set
Recall that the likelihood is a product over the games: \[ \begin{aligned} & L(\theta) = \prod_{i=1}^{N} P\left( y^{(i)} \,;\, \theta \right) \, P\left( x^{(i)} \mid y^{(i)} \,;\, \theta \right) \end{aligned} \]
Try deriving the likelihood yourself in the exercise below.
Answer:
Using the exponent trick, the likelihood becomes the following expression, with the exponents acting as switches: \[ \begin{aligned} & L(\theta) = \prod_{i=1}^{N} \underbrace{\left[ \pi^{y^{(i)}} (1 - \pi)^{(1 - y^{(i)})} \right]}_{P\left( y^{(i)} \right)} \\ & \qquad \underbrace{ \left[ \mathcal{N}\left( x^{(i)} \,;\, \mu_0, \sigma_0^2 \right) \right]^{(1 - y^{(i)})} \left[ \mathcal{N}\left( x^{(i)} \,;\, \mu_1, \sigma_1^2 \right) \right]^{y^{(i)}}}_{P\left( x^{(i)} \mid y^{(i)} \right)} \end{aligned} \]
Step 2: Derive the Log-Likelihood
Next, try deriving the log-likelihood in the exercise below.
Answer:
Taking the log, the product becomes a sum, and each exponent comes down as a multiplier. The log-likelihood is: \[ \begin{aligned} \ell (\theta) &= \sum_{i = 1}^{N} \left[ y^{(i)} \log(\pi) + \left( 1 - y^{(i)} \right) \log(1 - \pi) \right] \\[0.5em] & \quad + \sum_{i = 1}^{N} \left( 1 - y^{(i)} \right) \log \mathcal{N}\left( x^{(i)} \,;\, \mu_0, \sigma_0^2 \right) \\[0.5em] & \quad + \sum_{i = 1}^{N} y^{(i)} \log \mathcal{N}\left( x^{(i)} \,;\, \mu_1, \sigma_1^2 \right) \end{aligned} \]
Steps 3 and 4: Take the Derivative and Solve
As before, the log-likelihood nicely decomposes into three terms: one for \(\pi\) (using every game), one for \(\mu_0\) and \(\sigma_0^2\) (using only the games our team lost), and one for \(\mu_1\) and \(\sigma_1^2\) (using only the games our team won). Once again, we can optimize each term separately.
The MLE of \(\pi\) is the same as in the previous section. \[ \begin{aligned} & \hat{\pi}_{\text{MLE}} = \frac{N_{\text{win}}}{N} \end{aligned} \]
We will not formally derive the MLEs of the mean and variance of the Gaussian distributions, since this will be the focus of the Gaussian Discriminant Analysis chapter. For now, please take on faith that the expressions below are correct. You can also check them against your intuition. They compute the sample mean and sample variance of the opponent ratings, separately for the games our team lost and the games our team won.
\[ \begin{aligned} & \hat{\mu}_{0, \text{MLE}} = \frac{1}{N_{\text{loss}}} \sum_{i : y^{(i)} = 0} x^{(i)}, \\ & \hat{\sigma}_{0, \text{MLE}}^2 = \frac{1}{N_{\text{loss}}} \sum_{i : y^{(i)} = 0} \left( x^{(i)} - \hat{\mu}_{0, \text{MLE}} \right)^2 \end{aligned} \]
\[ \begin{aligned} & \hat{\mu}_{1, \text{MLE}} = \frac{1}{N_{\text{win}}} \sum_{i : y^{(i)} = 1} x^{(i)}, \\ & \hat{\sigma}_{1, \text{MLE}}^2 = \frac{1}{N_{\text{win}}} \sum_{i : y^{(i)} = 1} \left( x^{(i)} - \hat{\mu}_{1, \text{MLE}} \right)^2 \end{aligned} \]
You can verify that these expressions give the same values we computed in the previous section once you plug in the observed data. \[ \begin{aligned} & \hat{\pi}_{\text{MLE}} = 0.5 \\ & \hat{\mu}_{0, \text{MLE}} = 3.75 \quad\quad \hat{\sigma}_{0, \text{MLE}}^2 = 0.5625 \\ & \hat{\mu}_{1, \text{MLE}} = 2 \quad\quad \hat{\sigma}_{1, \text{MLE}}^2 = 0.25 \end{aligned} \]
3.4 Inference
With the parameters estimated, we can now return to our original question:
What is the probability that our team wins against an opponent whose rating is \(3.5\)?
As in Section 2.4, no model parameter directly corresponds to this probability. We can follow the same steps as before to express it using the model parameters.
In this derivation, we abuse notation slightly. Since \(X\) is continuous, \(P(X = 3.5)\) and \(P(X = 3.5 \mid Y = y)\) are really probability densities, not probabilities. We write them as probabilities to keep the notation consistent with Section 2.4. Bayes’ rule, the law of total probability, and the product rule still hold when densities replace these probabilities. Try the derivation yourself in the exercise below.
Answer:
Using Bayes’ rule, the law of total probability, and the product rule, we can express the probability using the model parameters: \[ \begin{aligned} P\left( Y = 1 \mid X = 3.5 \right) & = \frac{P\left( Y = 1 \right) \, P\left( X = 3.5 \mid Y = 1 \right)}{P\left( X = 3.5 \right)} \\[0.5em] & = \frac{P\left( Y = 1 \right) \, P\left( X = 3.5 \mid Y = 1 \right)} {P\left( Y = 0 \right) \, P\left( X = 3.5 \mid Y = 0 \right) + P\left( Y = 1 \right) \, P\left( X = 3.5 \mid Y = 1 \right)} \\[0.5em] & = \frac{\pi \, \mathcal{N}\left( 3.5 \,;\, \mu_1, \sigma_1^2 \right)} {(1 - \pi) \, \mathcal{N}\left( 3.5 \,;\, \mu_0, \sigma_0^2 \right) + \pi \, \mathcal{N}\left( 3.5 \,;\, \mu_1, \sigma_1^2 \right)} \end{aligned} \]
This model has a name. A generative classifier with a Bernoulli class prior and Gaussian class-conditionals is called Gaussian Discriminant Analysis (GDA), which will be the topic of one of our future chapters.
4 Many Features: The Curse of Dimensionality
The single-feature models worked well. What happens when we use several features? Suppose we record four binary features for each game: whether the game was played at home (\(X_1\)), whether our team won its previous game (\(X_2\)), whether a key player was available (\(X_3\)), and whether the game was played on a weekend (\(X_4\)).
| \(X_1\) | \(X_2\) | \(X_3\) | \(X_4\) | \(Y\) (Win?) | Count |
|---|---|---|---|---|---|
| 1 | 1 | 1 | 1 | 1 | 3 |
| 1 | 1 | 1 | 1 | 0 | 4 |
| 1 | 1 | 1 | 0 | 1 | 0 |
| 1 | 1 | 1 | 0 | 0 | 2 |
| \(\vdots\) | \(\vdots\) | \(\vdots\) | \(\vdots\) | \(\vdots\) | \(\vdots\) |
The product rule factors a joint distribution over two variables. With five variables, we need the chain rule, which generalizes the product rule to any number of variables.
Definition: For random variables \(A_1, A_2, \ldots, A_n\), the chain rule states that \[ \begin{aligned} & P(A_1, A_2, \ldots, A_n) \\ & = P(A_1) \, P(A_2 \mid A_1) \, P(A_3 \mid A_1, A_2) \cdots P(A_n \mid A_1, \ldots, A_{n-1}) \end{aligned} \]
Using the chain rule, we can factor the joint distribution of a single game one variable at a time, starting with \(Y\):
\[ \begin{aligned} & P\left( y, x_1, x_2, x_3, x_4 \right) \\ & = P\left( y \right) \, P\left( x_1 \mid y \right) \, P\left( x_2 \mid x_1, y \right) \, P\left( x_3 \mid x_1, x_2, y \right) \, P\left( x_4 \mid x_1, x_2, x_3, y \right) \end{aligned} \]
Every factor is a (conditional) Bernoulli distribution, which we already know how to estimate by counting. The difficulty arises when we count the number of parameters required.
The simplest way to count the parameters is to consider the joint distribution directly. With four binary features and a binary outcome, we have five binary variables, so the joint distribution assigns a probability to each of \[ \begin{aligned} & 2^5 = 32 \end{aligned} \] combinations of values. Because these probabilities must sum to \(1\), we need \(31\) parameters to specify the joint distribution. More generally, with \(d\) binary features, we need \[ \begin{aligned} & 2^{d+1} - 1 \end{aligned} \] parameters.
The chain rule factorization requires the same number of parameters, since it is simply another way of writing the full joint distribution. However, the factorization reveals where these parameters accumulate. Each factor requires one parameter for every combination of the variables it conditions on, so the later factors require far more parameters than the earlier ones. For example, \(P(Y)\) requires only \(1\) parameter, while the last factor, which conditions on all four other variables, requires \(16\) on its own.
This leads to two problems. The first problem is parameter explosion. As shown above, the number of parameters grows exponentially with the number of features.
The second problem is sparse data. Most combinations of the conditioning variables have few or no examples. For example, among the wins where the first three features all equal \(1\), the table shows no games with \(X_4 = 0\). The MLE therefore assigns this outcome a probability of zero: \[ \begin{aligned} & P\left( X_4 = 0 \mid X_1 = 1, X_2 = 1, X_3 = 1, Y = 1 \right) = 0 \end{aligned} \]
The next page presents a solution to these problems.
5 Summary
We started this page with a question the previous page couldn’t answer. Instead of asking “What’s the probability of winning?”, we asked “What’s the probability of winning, given something we know about the game?” Adding that one feature turned a single Bernoulli parameter into a small system of them, fit by the same MLE machinery as before. The product rule let us build the joint distribution out of pieces we already knew how to handle, and the log-likelihood always decomposed into groups that could each be maximized independently. After learning the parameters, we needed to use Bayes’ rule to calculate the probability of the outcome given the feature. This same recipe worked even when the feature became continuous.
The moment we tried to extend this to several features, the parameter count exploded exponentially and pure counting became unreliable because most feature combinations had few or no examples. On the next page, we will focus on a simple method that tries to tackle these problems.