Stochastic Gradient Descent
Learning Objectives
After reading this page, you should be able to:
- Define full batch gradient descent, stochastic gradient descent, and mini-batch gradient descent.
- Explain the trade-off between computational cost and gradient accuracy across the three variants of gradient descent.
- Explain how the batch size and learning rate affect the behaviour of SGD.
- Explain how plateaus and ravines affect the behaviour of SGD.
1 Introduction
For both linear and logistic regression, we turned a learning problem into an optimization problem. In both cases, the hypothesis space is restricted to functions of the form \(y = f(\mathbf{w}^\top \mathbf{x})\), where \(f(z) = z\) for linear regression and \(f(z) = \sigma(z)\) for logistic regression. In both cases, we chose loss functions that (1) align with our intuition for which hypothesis is preferable, and (2) are compatible with gradient-based optimization. That is, they are differentiable and provide informative gradient signals. Although we did not motivate the probabilistic interpretations explicitly, both models make probabilistic assumptions about the data that more concretely justify the choice of loss functions.
In both situations, gradient descent is one of the strategies that can be used to solve the optimization problem. Although other optimization strategies are possible (e.g., direct solution), we focused on gradient descent since it is more general-purpose and can be used for other models.
This page discusses an optimization algorithm called stochastic gradient descent (SGD), which is a variant of gradient descent useful in settings where the training dataset is large. SGD is broadly applicable to many machine learning models, including linear regression, logistic regression, and neural networks.
2 Variants of Gradient Descent
2.1 (Full Batch) Gradient Descent
Recall from our discussion of linear and logistic regression that gradient descent is an iterative optimization algorithm. At each iteration, we compute the gradient of the cost function with respect to the model parameters and update the parameters by taking a step in the direction opposite to the gradient.
Recall that the cost function \(\mathcal{E}(\mathbf{w})\) is defined as the average loss over the entire training set.
\[\mathcal{E}(\mathbf{w}) = \frac{1}{N} \sum_{i=1}^N \mathcal{L}^{(i)}(y^{(i)}, t^{(i)})\]
The gradient descent update rule is \[\begin{align*} \mathbf{w} &\leftarrow \mathbf{w} - \alpha \nabla_{\mathbf{w}} \mathcal{E}(\mathbf{w}) \\ &= \mathbf{w} - \frac{\alpha}{N} \sum_{i=1}^N \nabla_{\mathbf{w}} \mathcal{L}^{(i)}(y^{(i)}, t^{(i)}) \end{align*}\]
where \(\alpha\) is the learning rate and \(N\) is the total number of training examples.
Definition: Full batch gradient descent computes the gradient using all the training examples at each iteration. The gradient descent update rule is \[\begin{align*} \mathbf{w} &\leftarrow \mathbf{w} - \frac{\alpha}{N} \sum_{i=1}^N \nabla_{\mathbf{w}} \mathcal{L}^{(i)}(y^{(i)}, t^{(i)}) \end{align*}\]
This approach has a significant computational limitation. At each iteration, we must compute the gradient for every single training example before we can update the parameters. For a training set with \(N\) examples, this requires \(N\) forward passes and \(N\) gradient computations per iteration. If \(N\) is large (say, millions of examples), this becomes computationally expensive and can make training prohibitively slow.
Moreover, full batch gradient descent requires storing all training data in memory simultaneously, which may be infeasible for very large datasets.
The solution is to estimate the gradient using only a subset of the data at each iteration. This leads us to stochastic gradient descent.
2.2 Stochastic Gradient Descent
Let’s explore an approach at the other extreme. What if we estimate the gradient using one random example? The algorithm would go as follows.
Choose an example \(i\) uniformly at random.
Perform the update using the gradient for this example. \[\mathbf{w} \leftarrow \mathbf{w} - \alpha \nabla_{\mathbf{w}} \mathcal{L}^{(i)}(y^{(i)}, t^{(i)})\]
Repeat.
Definition: Stochastic gradient descent (SGD) estimates the gradient using one randomly sampled training example \(i\) at each iteration. The gradient descent update rule is \[\begin{align*} \mathbf{w} &\leftarrow \mathbf{w} - \alpha \nabla_{\mathbf{w}} \mathcal{L}^{(i)}(y^{(i)}, t^{(i)}) \end{align*}\]
This approach has several benefits. The cost per update is very small, and it is independent of the training set size \(N\). Since each update needs only one example, this approach can also handle the online learning scenario, where data arrives over time. Mathematically, it provides an unbiased estimate of the full batch gradient, as long as we perform uniform random sampling.
However, it has some obvious problems. Although the gradient estimate is unbiased, it has extremely high variance. We are essentially optimizing a different function at each step. Also, because we are using only one example, we cannot take advantage of the optimizations for vectorization implemented by most packages.
2.3 Mini-Batch Gradient Descent
Let’s explore a compromise between the two approaches, called mini-batch gradient descent. Instead of computing the gradient \(\nabla_{\mathbf{w}} \mathcal{E}(\mathbf{w})\) by using all the training examples or one example only, we will compute the gradient using a mini-batch, a randomly chosen subset of training examples. Let \(\mathcal{B} = \{i_1, i_2, \ldots, i_k\}\) denote a mini-batch with \(k\) training examples.
Definition: A mini-batch (or simply batch) is a subset of training examples used to estimate the gradient at each iteration. The batch size \(k\) is the number of training examples used to estimate the gradient at each iteration.
Using a mini-batch, we can estimate the gradient \(\nabla_{\mathbf{w}} \mathcal{E}(\mathbf{w})\) as follows.
\[\begin{align*} \nabla_{\mathbf{w}} \mathcal{E}(\mathbf{w}) = \frac{1}{N} \sum_{i=1}^N \nabla_{\mathbf{w}} \mathcal{L}^{(i)}(y^{(i)}, t^{(i)}) \approx \frac{1}{k} \sum_{i \in \mathcal{B}} \nabla_{\mathbf{w}} \mathcal{L}^{(i)}(y^{(i)}, t^{(i)}) \end{align*}\]
Definition: Mini-batch gradient descent is a variant of gradient descent that estimates the gradient using a randomly sampled subset (mini-batch \(\mathcal{B}\)) of training examples at each iteration. Thus, the gradient descent update rule becomes \[\begin{align*} \mathbf{w} &\leftarrow \mathbf{w} - \alpha \frac{1}{k} \sum_{i \in \mathcal{B}} \nabla_{\mathbf{w}} \mathcal{L}^{(i)}(y^{(i)}, t^{(i)}) \end{align*}\]
A note on terminology. We introduced stochastic gradient descent as using one random example to estimate the gradient at each iteration. In practice, however, people often use the term SGD loosely to refer to mini-batch gradient descent instead. We adopt this convention for the remainder of the course. From now on, SGD refers to mini-batch gradient descent.
The term “stochastic” refers to the random sampling of mini-batches, which introduces randomness into the optimization process. How do we sample a mini-batch? In theory, we should sample the examples independently and uniformly with replacement. In practice, the sampling is performed using a much simpler approach. Every time we go through the training data, we randomly permute the training set, partition it into mini-batches, and then process the mini-batches one at a time. This procedure involves two important concepts: iterations and epochs.
Definition: An epoch is one complete pass through the entire training dataset. In SGD, one epoch consists of processing all mini-batches exactly once.
Definition: An iteration is a single update step in the optimization algorithm. In SGD, one iteration corresponds to processing one mini-batch and updating the parameters once.
For a dataset with \(N\) training examples, SGD with a batch size \(k\) can be summarized as follows.
Initialize the model parameters \(\mathbf{w}\).
For each epoch,
- Randomly permute the training set.
- Partition the permuted training set into \(\frac{N}{k}\) mini-batches, each of size \(k\), so that each data point appears in exactly one mini-batch.
- For each mini-batch \(\mathcal{B}\),
- Compute the gradient estimate using the examples in the mini-batch. \[\nabla_{\mathbf{w}} \mathcal{E}(\mathbf{w}) \approx \frac{1}{k} \sum_{i \in \mathcal{B}} \nabla_{\mathbf{w}} \mathcal{L}^{(i)}(y^{(i)}, t^{(i)})\]
- Update the parameters. \[\mathbf{w} \leftarrow \mathbf{w} - \alpha \frac{1}{k} \sum_{i \in \mathcal{B}} \nabla_{\mathbf{w}} \mathcal{L}^{(i)}(y^{(i)}, t^{(i)})\]
Continue until convergence or until a stopping criterion is met. For example, we may stop after a fixed number of epochs or when the cost function stops decreasing.
In the algorithm above, an iteration is one execution of the inner loop, and an epoch is one execution of the outer loop. Thus, each epoch consists of \(\frac{N}{k}\) iterations.
3 Visualizations, Hyperparameters, and Geometry
3.1 GD vs. SGD Visualization
To understand the difference between full batch gradient descent (GD) and stochastic gradient descent (SGD), it is helpful to visualize how each algorithm navigates the loss landscape. The figure below contrasts the optimization paths taken by full batch gradient descent and stochastic gradient descent on a regression problem with features \(x_1\) and \(x_2\). We will learn a model \(y = w_1 x_1 + w_2 x_2\). Note that this model does not contain a bias term (no \(b\) or \(w_0\)), so that the weight space is 2D.
Here is a sample of the training data used for this visualization.
This table shows only the first few rows of the dataset. The full training set contains 100 data points.
Both visualizations show the 2D weight space, and the contours show \(\mathcal{E}\), the average loss across the entire dataset. In GD, we compute \(\nabla_{\textbf{w}}\mathcal{E}\) by averaging \(\nabla_{\textbf{w}}\mathcal{L}\) across the entire dataset. The negative of this gradient points in the direction of steepest descent of \(\mathcal{E}\).
In contrast, SGD uses only a mini-batch to estimate the gradient at each iteration. This estimated gradient may not point exactly in the direction of steepest descent of the full training loss, because it is computing the steepest descent of a different optimization problem: minimizing the loss on the current mini-batch. As a result, the optimization path is more “noisy” or “jagged” compared to full batch gradient descent. The visualization shows two sets of contours. The full training loss contours remain constant throughout optimization. The mini-batch loss contours change with each iteration as different mini-batches are processed.
Step: 0 / 100 Initial (w_1): 0.5 Initial (w_2): 2.5
Full batch GD versus SGD paths
Answer:
The blue contours show the full training loss \(\mathcal{E}(\mathbf{w})\), which averages the loss over all \(N\) training examples. The training set does not change, so this function stays the same in every iteration. The orange contours show the average loss over the current mini-batch. Each iteration uses a different mini-batch, so this function changes in every iteration.
SGD follows the gradient of the mini-batch loss, not the gradient of the full training loss. Thus, the gradient estimate at any single iteration can point in a different direction from the true gradient. However, since the mini-batches are sampled at random, the gradient estimate is correct on average. It is a noisy but unbiased estimate of the true gradient.
3.2 Choosing SGD Hyperparameters
The batch size and learning rate are two key settings that control the behaviour of SGD. As with other settings, there are important trade-offs involved in choosing their values.
3.2.1 Batch Size
The batch size involves a trade-off between computational efficiency and gradient quality. If the batch size is too large, we run into the same computational issues as in full batch gradient descent. Each iteration becomes expensive because we must store and process many examples. However, larger batches provide more accurate gradient estimates, leading to smoother optimization paths that converge more reliably.
If the batch size is too small, the gradient estimate becomes noisy because it is computed from only a few examples. This noise can make the optimization path erratic and may slow down convergence. However, small batches have computational advantages. Each iteration is fast and requires less memory. Interestingly, the noise introduced by small batches can sometimes help the optimizer escape poor local minima, as we will discuss later.
Batch sizes are commonly chosen to be powers of 2, such as \(k = 32\) to \(k = 256\), though the optimal choice depends on the problem, the dataset size, and available computational resources. For certain types of data, such as videos, \(k=1\) may be necessary due to memory limits.
3.2.2 Learning Rate
The learning rate \(\alpha\) controls the size of each parameter update. If the learning rate is too small, the weight updates are tiny in each iteration, so it can take a very long time for the weights to get close to optimal values.
If the learning rate is too large, the weight updates may overshoot the optimal values, causing the optimization to diverge or oscillate around the minimum without converging. In full batch gradient descent, a learning rate that is too large will cause the cost function to increase, leading to clear divergence. However, in SGD, the situation is more nuanced. Because the gradient estimate is noisy, even a moderately large learning rate may not cause immediate divergence, but it can still prevent the algorithm from converging to a good solution.
3.2.3 Interactions
The learning rate and batch size interact in important ways. When using smaller batches (which produce noisier gradient estimates), it is often beneficial to use a smaller learning rate to compensate for the increased noise. Conversely, larger batches with more accurate gradient estimates can often tolerate larger learning rates.
The figure below illustrates how different combinations of batch size and learning rate affect the optimization path. By comparing side-by-side visualizations, we can see how these hyperparameters interact to produce different convergence behaviours.
Batch size and learning rate effects
Answer:
SGD takes one gradient descent step per mini-batch. With \(N\) training examples and a batch size of \(k\), each epoch contains \(\frac{N}{k}\) mini-batches, so it takes \(\frac{N}{k}\) steps. Over \(E\) epochs, SGD takes \(E \cdot \frac{N}{k}\) steps in total. If \(E\) is fixed and \(k\) increases, then the number of steps decreases. For example, with \(N = 100\), a batch size of \(10\) gives \(10\) steps per epoch, while a batch size of \(50\) gives only \(2\).
3.3 The Geometry of GD and SGD
Recall from Fundamental Idea #4 that machine learning describes geometric processes. With this perspective, we can think of models as objects with geometry. Some models are more similar to each other than to others. Optimization algorithms like gradient descent and SGD make small, incremental changes to models, moving through the space of possible models.
This geometric perspective helps us understand where optimization algorithms can struggle. The loss landscape (the function that maps model parameters to loss values) can have challenging geometric structures that make optimization difficult. In this section, we explore some of these challenges.
3.3.1 Plateaus
We have already encountered examples of plateaus in our discussion of logistic regression, though we didn’t call them by that name at the time. A plateau is a region of the loss landscape where the gradient is very small (or zero) across a large area. On a plateau, the loss function is nearly flat, meaning that small changes to the parameters result in almost no change to the loss. This creates a problem for gradient-based optimization. If the gradient is near zero, the parameter updates will be tiny, and the algorithm will make very slow progress.
Definition: A plateau is a region in the loss landscape \(\mathcal{E}\) where the gradient \(\nabla_{\textbf{w}}\mathcal{E}\) is small, despite the cost \(\mathcal{E}(\textbf{w})\) being high.
We saw an example of a plateau in our (failed) attempt to build a linear classification model, when we paired the squared error loss with the sigmoid activation function. With this approach, we saw that when the model makes predictions with very high confidence (\(y \approx 0\) or \(y \approx 1\)), the derivative \(\frac{dy}{dz}\) becomes very small. This happens even when the model is confident and wrong (\(t = 1\) when \(y \approx 0\) or vice versa). Thus, even when the model is making incorrect predictions, the gradient signal is weak, and learning is slow.
From a geometric perspective, plateaus represent “flat valleys” in the loss landscape. The optimizer can get stuck in these regions, making minimal progress even though there may be better solutions nearby. The randomness introduced by SGD can sometimes help escape plateaus, as the noisy gradient estimates may occasionally point in directions that lead out of the flat region. However, it is best to construct loss functions to avoid plateaus if possible.
3.3.2 Ravines
A ravine is a region in the loss landscape where the cost function changes much more rapidly in one direction than in another. This creates a long, narrow valley that can make optimization challenging, as the gradient descent algorithm may make rapid progress along one dimension while making very slow progress along another.
Definition: A ravine is a region in the loss landscape \(\mathcal{E}(\textbf{w})\) where \(\mathcal{E}\) changes much more rapidly along one direction in the weight space than along a perpendicular direction. In 2D/3D weight space, the loss landscape can be visualized as an elongated valley that can slow down convergence when using GD/SGD.
We first encountered ravines in Feature Scaling for Gradient Descent, where features on very different scales caused gradient descent to zig-zag across a narrow valley. Here, we revisit that example with SGD and look more closely at why a single learning rate struggles.
Suppose we use the following dataset to fit the linear regression model \(y = x_1 w_1 + x_2 w_2\), except that \(x_1\) is multiplied by a large number. For example, one feature can be 100 or more times the scale of another. This could happen if \(x_1\) represents house area in square feet (ranging from 1000 to 5000) while \(x_2\) represents lot size in acres (ranging from 0.1 to 2.0).
So why is this a problem? Let’s start by thinking about the magnitudes of the optimal weights.
Answer:
We would expect \(w_2\) to be larger in magnitude than \(w_1\) at the optimal solution. Since \(x_1\) is large, a small \(w_1\) is enough for \(w_1 x_1\) to contribute significantly to \(y\). Since \(x_2\) is small, \(w_2\) must be large for \(w_2 x_2\) to make a meaningful contribution to \(y\).
Suppose we initialize \(w_1\) and \(w_2\) in the usual way, such as to values close to 0. Will \(w_1\) and \(w_2\) be able to reach their optimal values? Recall from Feature Scaling for Gradient Descent that \(\frac{\partial \mathcal{L}}{\partial w_j} = x_j(y-t)\), so the partial derivative for each weight is scaled by its feature.
But this is not what we want! Even though we want \(w_2\) to be larger, the gradient \(\frac{\partial \mathcal{L}}{\partial w_2} = x_2(y-t)\) is much smaller than \(\frac{\partial \mathcal{L}}{\partial w_1} = x_1(y-t)\). This means \(w_1\) receives much larger updates than \(w_2\), even though \(w_2\) needs to make more progress to reach its optimal (larger) value!
The visualization below shows exactly this problem. Note that to make it easy to see what is going on, we have only scaled \(x_1\) by a factor of 3 (not 100). A 100-fold increase in \(x_1\) actually makes a much more drastic difference, and the figure vastly understates the narrowness of the ravine! In reality, the ravine would be much more elongated, making the optimization challenge even more pronounced.
Ravine from feature scaling
In this visualization, we see that we face a conundrum. If the learning rate is too small (left panel, \(\alpha = 0.02\)), then progress in the \(w_2\) direction is extremely slow, even though the algorithm quickly reaches the minimum in the \(w_1\) direction. However, if we increase the learning rate to speed up progress in the \(w_2\) direction (right panel, \(\alpha = 0.04\)), the optimization path begins to oscillate back and forth in the \(w_1\) direction while still making no meaningful progress in the \(w_2\) direction.
The fundamental issue is that we are using a single learning rate for both dimensions, when in reality we need different learning rates along different directions. The ideal solution would be to use a larger learning rate for \(w_2\) (the flat direction) and a smaller learning rate for \(w_1\) (the steep direction).
We also saw in Feature Scaling for Gradient Descent that highly correlated features create ravines as well, and these ravines are not axis-aligned. While standardization helps address ravines caused by features with different scales, it cannot remove ravines caused by highly correlated features. Figure 4 shows such a ravine in the SGD setting.
Thus, while SGD is a powerful and widely used optimization algorithm, it is often not sufficient on its own. Advanced optimization methods like momentum, Adam, and other adaptive optimizers attempt to address these challenges by using second-order information and per-dimension learning rate adaptation. However, a detailed discussion of these methods is beyond the scope of these notes.
3.3.3 Other Landscapes
The loss landscapes we’ve explored, plateaus and ravines, represent just a few examples of the complex optimization surfaces encountered in machine learning. Real-world loss landscapes can exhibit many other pathological structures, such as multiple local minima and saddle points. A comprehensive treatment of these phenomena is beyond the scope of these materials. We hope that understanding the basic challenges helps explain why optimization in machine learning remains an important and active area of research.
4 Summary
This chapter compared three ways to estimate the gradient in gradient descent. Full batch gradient descent computes the exact gradient using all \(N\) training examples, but each update becomes expensive in time and memory when \(N\) is large. Stochastic gradient descent sits at the other extreme. It estimates the gradient using one random example, so each update is cheap and independent of \(N\). However, its gradient estimates have high variance, and it cannot take advantage of vectorization. Mini-batch gradient descent is a compromise between the two. It estimates the gradient using a small random subset of examples, and it is what people usually mean by SGD in practice. Every time we go through the training data, we reshuffle it and process the mini-batches one at a time. Each update is an iteration, and each full pass over the training data is an epoch.
The batch size and the learning rate are the two key hyperparameters of SGD. A larger batch gives more accurate gradient estimates but makes each iteration more expensive. A smaller learning rate makes slow progress, while a larger one may overshoot the minimum. The two also interact, so they should be tuned together.
This chapter illustrates Fundamental Idea #1 (learning is optimization). Once we frame learning as minimizing a cost function, we can swap in a different optimizer without changing the model or the loss. SGD applies to linear regression, logistic regression, and many other models. However, these choices are not fully independent. The model, the loss, and the data together shape the loss landscape that the optimizer must navigate.
This is where Fundamental Idea #4 (ML describes geometric processes) comes in. Gradient descent and SGD move through the weight space in small steps, so the geometry of the loss landscape determines how well they work. On a plateau, the gradient is small even though the cost is high, so progress is slow. In a ravine, the cost changes much faster in one direction than another, so a single learning rate cannot suit both directions. Ravines can come from features on different scales, which standardization helps with, and from correlated features, which standardization cannot fix. More advanced optimizers, such as momentum and Adam, address these challenges, but they are beyond the scope of these notes.
Finally, SGD reflects Fundamental Idea #5 (ML demands a probabilistic lens). Randomly sampling mini-batches makes the gradient estimate noisy but unbiased. In this chapter, we treated this noise as an obstacle to manage through the batch size and learning rate. In future study, we will see that this randomness can also be beneficial. The noisy gradient estimates can help SGD escape shallow local minima that might trap full batch gradient descent, allowing it to explore the loss landscape more broadly and potentially find better solutions.