Limitations of Linear Models

Learning Objectives

After reading this page, you should be able to:

  1. Define linear separability.
  2. Prove that the XOR dataset is not linearly separable.
  3. Explain how adding a feature can make a dataset linearly separable.
  4. Describe the limitations of hand-crafting features.

1 Introduction

Linear models such as linear regression, logistic regression, and softmax regression are foundational tools in machine learning. They are efficient, interpretable, and often perform surprisingly well. However, they all share a common structural assumption: the model’s prediction is based on a linear function of the input features.

For binary classification, this assumption imposes a fundamental restriction. Linear classifiers can only learn linear decision boundaries, which are hyperplanes of the form \(\{\mathbf{x} : \mathbf{w}^\top \mathbf{x} + b = 0\}\). When the underlying class structure is not linearly separable, this assumption can become a major limitation.

This limitation explains why we need more advanced models, such as neural networks, which we will see in the next chapter. This section reasons about the limitations of linear models in detail. We will explore ways around this limitation through feature mapping and explain why it is not always sufficient.

2 Linear Models and Linear Separability

As we saw in the previous sections, a logistic regression classifier computes predictions as \[y = \sigma(\mathbf{w}^\top \mathbf{x} + b),\] where \(\sigma(z) = \frac{1}{1+e^{-z}}\) is the sigmoid function, \(\mathbf{w}\) is a weight vector, and \(b\) is a bias term. If we use \(\sigma(z) = 0.5\) as the decision boundary threshold, then the decision boundary is where \(z = 0\). This boundary is the set \[\{\mathbf{x} : \mathbf{w}^\top \mathbf{x} + b = 0\},\] which is a hyperplane.

For binary classification, this means we can only separate the two classes if they are linearly separable.

Definition: Two sets of points are linearly separable if there exists a hyperplane that perfectly separates them.

Linear vs non-linear separability

Figure 1: (Left) The two classes can be perfectly separated by a single straight line (dashed). (Right) The ring pattern cannot be separated by a straight line. The inner class is blue and the outer class is orange.

If the true decision boundary is curved or has multiple disconnected regions, a linear model will not be able to perfectly classify this data. This is true no matter how many training data points we have and no matter what optimization method we use.

Consider the leaf classification problem we have used throughout this text. The visualization below shows both the data space (left) and the weight space (right). Try selecting weights in the right panel of Figure 2 to find a decision boundary that perfectly separates the two classes.

Leaf weight-space explorer

Figure 2: Interactive. With \(b = -150\) fixed, the data space is shown on the left and the weight space on the right. Click anywhere in the weight space to choose a pair \((w_1, w_2)\) and watch the decision boundary update in the data space. Can you find a boundary that correctly separates all oak leaves from all maple leaves?

You should see that, no matter your choice of \((w_1, w_2)\), it is impossible to achieve 100% training accuracy with a linear decision boundary. However, we kept \(b\) fixed at \(-150\). How can we know for sure that the dataset is really not linearly separable?

3 The Mathematics of Linear Separability

In this section, we will consider a dataset for a simple logical function and prove that it is not linearly separable. Our goal is to model the exclusive OR (XOR) function for binary inputs \(0\) and \(1\). XOR outputs 1 when the two inputs are different, and 0 otherwise. XOR can be defined with four data points, as shown in Table 1 below.

Table 1: The four XOR data points.
\(x_1\) \(x_2\) \(t\)
0 0 0
0 1 1
1 0 1
1 1 0

The XOR function is more than a convenient toy example for an introductory course. It also played a famous role in the history of AI. The story begins in 1958, when Frank Rosenblatt unveiled the perceptron, a linear classifier that learned from examples. The excitement was enormous. The New York Times reported that the Navy expected the perceptron to “walk, talk, see, write, reproduce itself and be conscious of its existence.” Eleven years later, however, Marvin Minsky and Seymour Papert published a book titled Perceptrons. In it, they proved that a single-layer perceptron cannot compute even the simple XOR function. The result itself was a narrow technical claim about a specific architecture. However, it landed during a period of widespread, often inflated claims about what perceptrons would soon achieve, and many readers, including funding agencies, concluded that neural networks in general were a dead end. Networks with more layers could compute XOR, but at the time, no one knew how to train them. As a result, interest and funding for neural networks dried up, and the field’s decline became part of the first AI winter of the 1970s. Neural networks did not return to the spotlight until the 1980s.

The irony is that the limitation Minsky and Papert identified was easily fixable. Adding even a single hidden layer with a non-linear activation function turns XOR into a problem the network can solve perfectly. The XOR example is therefore worth remembering not just as a small classification problem, but as the moment that motivated the move from single-layer to multi-layer networks and the algorithm that makes them trainable.

Returning to the XOR example, the visualization below shows the four XOR data points in the feature space alongside an interactive weight space (with \(b = 1\) fixed). Try clicking in the weight space to explore different linear boundaries.

XOR weight-space explorer

Figure 3: Interactive. With \(b = 1\) fixed, the feature space is shown on the left and the weight space on the right. Click anywhere in the weight space to choose \((w_1, w_2)\) and see the resulting decision boundary. Can you find weights that correctly classify all four points?

Again, you should see that, no matter your choice of the two weights, it is impossible to achieve 100% training accuracy with a linear decision boundary. The following theorem states this result formally.

Theorem (XOR is not linearly separable): No linear classifier can correctly classify all four data points of the XOR function.

Let’s prove this theorem in two ways, first algebraically and then geometrically. The two proof strategies use similar ideas but express them differently.

3.1 Algebraic Perspective: Feasible Regions in the Weight Space

We saw in an earlier chapter that correctly classifying a training example places a constraint on the weight vector. Let’s apply this idea to logistic regression.

Consider a logistic regression classifier with a threshold cutoff of \(0.5\). For a positive example (\(t=1\)), we want the probability to be at least \(0.5\). This translates to the following constraint on \(b, w_1, w_2\). \[b + w_1 x_1 + w_2 x_2 \geq 0\] Similarly, for a negative example (\(t=0\)), the constraint on the weights is \[b + w_1 x_1 + w_2 x_2 < 0.\]

If a hypothesis correctly classifies all data points, the weights must satisfy all the constraints imposed by the data points. Table 2 lists the four constraints for the XOR problem.

Table 2: Constraints on \((b, w_1, w_2)\) imposed by each XOR example.
\(x_1\) \(x_2\) \(t\) Constraint
0 0 0 \(b < 0\)
0 1 1 \(b + w_2 \geq 0\)
1 0 1 \(b + w_1 \geq 0\)
1 1 0 \(b + w_1 + w_2 < 0\)

Question: Use an algebraic approach to prove that the four constraints cannot be satisfied at the same time.

Answer:

Adding the constraints from the two positive examples gives \[2b + w_1 + w_2 \geq 0.\] Adding the constraints from the two negative examples gives \[2b + w_1 + w_2 < 0.\] These inequalities are contradictory, so the feasible region is empty. No weights exist that correctly classify all four XOR points.

Figure 3 confirms this result. No matter where you click in the weight space, at least one point remains misclassified.

3.2 Geometric Perspective: Convexity in Feature Space

There is an equivalent geometric proof that works directly in the feature space. Figure 4 labels the four XOR points. The positive examples are \(A = (0, 1)\) and \(C = (1, 0)\). The negative examples are \(B = (0, 0)\) and \(D = (1, 1)\). The segment \(\overline{AC}\) connecting the two positive examples and the segment \(\overline{BD}\) connecting the two negative examples intersect at the point \(F = (0.5, 0.5)\).

XOR midpoint contradiction

Figure 4: The positive examples are \(A\) and \(C\), and the negative examples are \(B\) and \(D\). The segment \(\overline{AC}\) connects the two positive examples, and the segment \(\overline{BD}\) connects the two negative examples. The two segments intersect at \(F = (0.5, 0.5)\).

The proof relies on the fact that the decision boundary of a linear model separates the data space \(\mathbb{R}^d\) into two half-spaces.

Definition: A half-space in \(\mathbb{R}^d\) is the set of all points on one side of a hyperplane, \[H = \{\mathbf{x} \in \mathbb{R}^d : \mathbf{a}^\top \mathbf{x} + b \geq 0\}\] for some \(\mathbf{a} \in \mathbb{R}^d\) and \(b \in \mathbb{R}\). The boundary \(\{\mathbf{x} : \mathbf{a}^\top \mathbf{x} + b = 0\}\) is the hyperplane itself.

The proof is a proof by contradiction.

Question: Here is an outline of the proof with one step missing.

  1. Assume the XOR data is linearly separable.
  2. The decision boundary of a linear model separates the inputs into a positive half-space and a negative half-space.
  3. Each half-space is a convex set.
  4. ____________________
  5. The segment \(\overline{AC}\) lies in the positive half-space. So does \(F\).
  6. The segment \(\overline{BD}\) lies in the negative half-space. So does \(F\).
  7. \(F\) lies in both the positive and the negative half-spaces. This is a contradiction.

Step 5 concludes that the whole segment \(\overline{AC}\) lies in the positive half-space. However, we only know that the two endpoints \(A\) and \(C\) lie in the positive half-space. What property of convex sets makes step 5 valid? Write this property as step 4.

Answer:

For any two points in a convex set, the line segment connecting the two points also lies in the same convex set.

The property in step 4 is exactly what it means for a set to be convex. The definition below states this property formally. A point on the segment between \(\mathbf{x}^{(a)}\) and \(\mathbf{x}^{(b)}\) can be written as \(\lambda\mathbf{x}^{(a)} + (1-\lambda)\mathbf{x}^{(b)}\) for some \(\lambda \in [0,1]\).

Definition: A set \(S \subseteq \mathbb{R}^d\) is convex if for every two points \(\mathbf{x}^{(a)}, \mathbf{x}^{(b)} \in S\) and every \(\lambda \in [0,1]\), the point \(\lambda\mathbf{x}^{(a)} + (1-\lambda)\mathbf{x}^{(b)}\) also lies in \(S\). Intuitively, the line segment connecting any two points in \(S\) stays inside \(S\).

Step 3 claims that every half-space is convex. The next exercise asks you to verify this claim.

Question: Show that a half-space \(H\) defined below is convex. \[H = \{\mathbf{x} : \mathbf{a}^\top \mathbf{x} + b \geq 0\}\]

Answer: For any two points \(\mathbf{x}^{(a)}, \mathbf{x}^{(b)} \in H\), we have \[\mathbf{a}^\top \mathbf{x}^{(a)} + b \geq 0, \qquad \mathbf{a}^\top \mathbf{x}^{(b)} + b \geq 0.\] We can show that any convex combination of the two points also lies in this half-space. For any \(\lambda \in [0,1]\), \[ \begin{aligned} & \mathbf{a}^\top\!\bigl(\lambda\mathbf{x}^{(a)} + (1-\lambda)\mathbf{x}^{(b)}\bigr) + b \\ & = \lambda\bigl(\mathbf{a}^\top \mathbf{x}^{(a)} + b\bigr) + (1-\lambda)\bigl(\mathbf{a}^\top \mathbf{x}^{(b)} + b\bigr) \geq 0 \\ & \Rightarrow \lambda\mathbf{x}^{(a)} + (1-\lambda)\mathbf{x}^{(b)} \in H. \end{aligned} \]

We can now write out the full proof.

Proof. Suppose a linear classifier \[f(\mathbf{x}) = b + w_1 x_1 + w_2 x_2\] correctly classifies all four XOR points. Then the positive examples lie in the half-space \[H^+ = \{\mathbf{x} : f(\mathbf{x}) \geq 0\},\] and the negative examples lie in the half-space \[H^- = \{\mathbf{x} : f(\mathbf{x}) < 0\}.\] Both half-spaces are convex.

The point \(F\) is the midpoint of \(\overline{AC}\), since \[F = \frac{A + C}{2} = \frac{\begin{bmatrix}0\\1\end{bmatrix}+\begin{bmatrix}1\\0\end{bmatrix}}{2} = \begin{bmatrix}0.5\\0.5\end{bmatrix}.\] The point \(F\) is also the midpoint of \(\overline{BD}\), since \[F = \frac{B + D}{2} = \frac{\begin{bmatrix}0\\0\end{bmatrix}+\begin{bmatrix}1\\1\end{bmatrix}}{2} = \begin{bmatrix}0.5\\0.5\end{bmatrix}.\]

Since \(A\) and \(C\) lie in \(H^+\) and \(H^+\) is convex, \(F\) is also in \(H^+\).

Since \(B\) and \(D\) lie in \(H^-\) and \(H^-\) is convex, \(F\) is also in \(H^-\).

But \(H^+\) and \(H^-\) are disjoint, and \(F\) cannot be in both, a contradiction. Therefore, the XOR data is not linearly separable. \(\square\)

The algebraic and geometric proofs are closely related. In the algebraic proof, adding the constraints for the two positive examples shows that \(F\) must be classified as positive. Adding the constraints for the two negative examples shows that \(F\) must be classified as negative. In other words, adding two constraints is the same as evaluating \(f\) at the midpoint of the two points. Both proofs arrive at the same contradiction at the same point \(F\).

4 Feature Engineering as a Partial Solution

We have seen so far that linear models cannot capture datasets that are not linearly separable, even ones as simple as XOR. Sometimes, we can overcome this limitation by using feature maps.

Returning to our XOR problem, let’s see how we can use a linear model to perfectly classify the XOR data points by adding a new feature \(x_3 = x_1 x_2\). Table 3 shows the four points in the extended space \((x_1, x_2, x_3)\).

Table 3: The XOR dataset extended with the feature \(x_3 = x_1 x_2\).
\(x_1\) \(x_2\) \(x_3 = x_1 x_2\) \(t\)
0 0 0 0
0 1 0 1
1 0 0 1
1 1 1 0

Before we attempt to solve for a set of weights for our linear model, let’s try to get some intuition about why the new feature helps.

Question: Figure 5 shows the extended space \((x_1, x_2, x_3)\). Draw the four data points \(A\), \(B\), \(C\), and \(D\) in this space. Then draw a plane that separates the positive examples from the negative examples.

The extended feature space

Figure 5: An empty plot of the extended space \((x_1, x_2, x_3)\). The grey square is the original feature space at \(x_3 = 0\).

Answer:

Figure 6 shows the four points and one separating plane. The feature \(x_3\) is 0 for the first three points and 1 for \(D\). So the new feature lifts \(D\) above the other three points. The positive examples \(A\) and \(C\) are no longer stuck between the negative examples \(B\) and \(D\). A plane can now pass between the two classes.

XOR in the extended feature space

Figure 6: The four XOR points in the extended space \((x_1, x_2, x_3)\). The grey square is the original feature space at \(x_3 = 0\). The feature \(x_3 = x_1 x_2\) lifts \(D\) to \(x_3 = 1\). The shaded plane is one example of a plane that separates the positive examples \(A\) and \(C\) from the negative examples \(B\) and \(D\).

Now that we have gained some intuition about why the new feature helps, let’s try to derive the weights and bias for the linear model algebraically. Since this algebraic derivation is not the focus of this page, the following exercise also gives you the option to verify that a given set of weights and bias correctly classifies all four points.

Question: Find values of \((b, w_1, w_2, w_3)\) that correctly classify all four points.

Alternatively, verify that the following set of weights and bias correctly classifies all four points. \[b = -0.5, w_1 = 1, w_2 = 1, w_3 = -2\]

Answer:

The four constraints are \[\begin{align*} b &< 0 \\ b + w_2 &\geq 0\\ b + w_1 &\geq 0\\ b + w_1 + w_2 + w_3 &< 0 \end{align*}\] One solution is \[b = -0.5,\ w_1 = 1,\ w_2 = 1,\ w_3 = -2,\] which gives \[\begin{align*} f(0,0,0) &= -0.5 < 0\\ f(0,1,0) &= 0.5 > 0\\ f(1,0,0) &= 0.5 > 0\\ f(1,1,1) &= -0.5 < 0. \end{align*}\]

The exercise above gives the separating hyperplane \(x_1 + x_2 - 2\,x_3 - 0.5 = 0\) in the extended space. Substituting \(x_3 = x_1 x_2\) gives the decision boundary \(x_1 + x_2 - 2\,x_1 x_2 - 0.5 = 0\) in the original data space \((x_1, x_2)\). This boundary factors as \((2x_1 - 1)(2x_2 - 1) = 0\), so it consists of the two lines \(x_1 = 0.5\) and \(x_2 = 0.5\). Figure 7 visualizes this boundary, which is non-linear in the original data space.

Quadratic boundary for XOR

Figure 7: The XOR points are shown in the original \((x_1, x_2)\) space with the dashed boundary \(x_1 + x_2 - 2\,x_1 x_2 - 0.5 = 0\) induced by the feature map \(x_3 = x_1 x_2\). The boundary is the pair of lines \(x_1 = 0.5\) and \(x_2 = 0.5\). Unlike a single linear boundary, it correctly separates the two classes.

Other weights can give very different boundaries. For example, the weights \(b = -0.5\), \(w_1 = 1\), \(w_2 = 1\), and \(w_3 = -3\) also classify all four points correctly. Their decision boundary in the original data space is \(x_1 + x_2 - 3\,x_1 x_2 - 0.5 = 0\). This boundary is a hyperbola, as shown in Figure 8.

Hyperbolic boundary for XOR

Figure 8: The XOR points in the original \((x_1, x_2)\) space with the dashed boundary \(x_1 + x_2 - 3\,x_1 x_2 - 0.5 = 0\). The boundary is a hyperbola with two branches. It also separates the two classes correctly.

This feature engineering approach can create more expressive models, but the choice of features requires domain knowledge and manual experimentation. The right interactions are not obvious. Moreover, the number of possible features grows combinatorially. For \(d\) inputs, all degree-2 products give \(O(d^2)\) new features, which becomes infeasible for large \(d\). Adding features aggressively without regularization can also lead to overfitting.

5 Motivation for Neural Networks

This limitation of linear models motivates the next chapter. Rather than hand-crafting feature maps, we will build models that learn non-linear representations directly from data.

In the next chapter, we will motivate and consider neural networks from different perspectives. One perspective that you should remember is that the last step of a neural network is a linear model. The difference is that its features are learned rather than hand-set (Figure 9).

Learned features plus a linear classifier

Figure 9: An opaque learned transformation (hidden layers) maps the input \(\mathbf{x}\) to learned features \(\phi(\mathbf{x})\), followed by a linear classifier \(\mathbf{w}^\top \phi(\mathbf{x}) + b\).