Introducing the Neuron

Learning Objectives

After reading this page, you should be able to:

  1. Describe the main components of a biological neuron and how signals flow through it.
  2. Construct the mathematical model of an artificial neuron, \(y = f(\mathbf{w}^\top \mathbf{x} + b)\).
  3. Interpret logistic regression as a single artificial neuron.
  4. Interpret softmax regression as a one-layer neural network.
  5. Distinguish the different conventions used to draw neural network diagrams.

1 Introduction

The models we have already studied (logistic regression and softmax regression) turn out to be limited in a fundamental way. Because their decision boundary is always a straight line (or a flat hyperplane in higher dimensions), they cannot classify data that is not linearly separable. As we saw in Limitations of Linear Models, even a task as simple as exclusive OR defeats such a model. This limitation is what motivates more expressive models, and chief among them are neural networks.

Neural networks are hyped for a reason, but the mathematics behind them involves little more than multiplication, addition, and a few logarithms. What makes them powerful is scale: a network performs an enormous number of these simple operations, which is exactly why we lean on linear algebra to keep the bookkeeping manageable. In fact, one of the first things we will show is that the models we have already studied are the starting point: logistic and softmax regression can each be interpreted as a single artificial neuron, or equivalently a small one-layer neural network that applies a linear transformation to its input and then passes the result through an activation function. The natural fix for their limitation is then to stack more layers: by composing several linear transformations with nonlinear activation functions, a network can represent the curved, complex decision boundaries that a single layer never could.

This page opens with a brief look at biological neurons, which inspired the mathematical abstraction used throughout this chapter, and then shows how logistic and softmax regression can each be interpreted as a single-layer neural network.

After gaining familiarity with working with single neurons, we will build up to multi-layer perceptrons where many neurons are used, including in hidden layers where a neuron’s activity is not prescribed a specific meaning. From there, we consider how neural networks are trained with backpropagation. We will conclude this chapter with a discussion of the expressiveness of neural networks and how it resolves the limitations of linear models we saw earlier.

The structure of this section draws inspiration from Stanford’s CS231n materials.

2 Biological Neurons

The ideas behind neural networks come from neuroscience and biology, and are inspired by the brain. We therefore start our discussion of neural networks by discussing the brain. This section on biology is not critical and will be brief. However, it is interesting to make a connection between computer science and biology and to understand where some neural network terminology comes from.

Let’s start with the structure of a biological neuron. Figure 1 shows the main components of a neuron and how information flows through these components. First, the dendrites are the neuron’s input channels. Each dendrite receives signals from another neuron at a synapse or synaptic connection. The dendrites then carry those signals toward the cell body, also called the soma. Then, the cell body integrates all of these incoming signals. You can think of this integration as collecting evidence from many sources and combining the evidence into a single overall signal. If this combined signal is strong enough, the neuron fires: it sends an electrical impulse onward through the axon. This axon has connections to the dendrites of the next cell through other synapses. The axon thus acts as the neuron’s output channel.

Biological neuron anatomy

Figure 1: Interactive Dendrites are inputs, the cell body integrates, the axon outputs, and synapses connect cells. The central neuron receives signals from axons of neighbouring cells across synapses onto dendrites, integrates incoming signals at the cell body, and (when firing occurs) sends the signal through the axon to dendrites of downstream cells. Hover (or tab to focus) on a label to show the corresponding neuron component; click to pin the highlight.

Synapses are the connections between neurons, and they are especially important because they have different strengths. Some connections transmit signals more strongly than others, and these strengths can change over time. In particular, when one neuron repeatedly takes part in firing another, the synapse between them strengthens so that the first neuron more easily triggers the second in the future, a principle known as Hebbian learning (Hebb, 1949). At the chemical level, this signalling happens through neurotransmitters, molecules released across the synapse whose type and quantity determine how strongly the signal is passed on.

3 The Artificial Neuron

We now propose a mathematical model of a neuron that is far simpler than a real one. However, the model captures the important ideas about how a neuron processes information: inputs are combined, a decision is made about whether to “fire,” and the strengths of connections play a crucial role in learning. We are not aiming for biological accuracy, but for a clean mathematical abstraction that keeps only these important ideas. The diagram below shows this simplified model.

Artificial neuron model

Figure 2: Interactive Inputs \(x_i\), synaptic weights \(w_i\), bias \(b\), activation \(f\), and output. Hover (or tab to focus) on a label to show the corresponding neuron component; click to pin the highlight.

You can still recognize the main components we discussed earlier: axons deliver signals, synapses connect neurons, dendrites receive inputs, and the cell body integrates information. What’s new here is that each part now has a precise mathematical interpretation.

Figure 2 can be interpreted in two complementary ways. The first is a biological perspective, in which each part of the model corresponds to a part of a real neuron and the diagram shows how signals flow from inputs to output. The values \(x_1, x_2, x_3\) represent electrical signals arriving from other neurons via their axons. These signals pass through synapses into the dendrites, and each synapse has an associated weight \(w_1, w_2, w_3\) that models the strength and sign of that synaptic connection: a larger weight amplifies the signal, a smaller weight dampens it, and a negative weight produces an inhibitory effect. Inside the cell body, each incoming signal is scaled by its synaptic weight, producing terms \(w_1 x_1\), \(w_2 x_2\), and \(w_3 x_3\), which are then summed together. A bias term \(b\) captures the intrinsic excitability of the neuron, its tendency to fire independent of external input. The cell body therefore accumulates a total signal \(\sum_{j=1}^D w_j x_j + b\). This total is passed through a non-linear activation function \(f\) that determines whether and how the neuron fires given the signal. The output of the neuron can be written \(y = f(\sum_j w_j x_j + b)\), and this output signal is passed along the axon to downstream neurons.

Next, Figure 2 can also be interpreted from a machine learning perspective. From this perspective, \(x_1, x_2, x_3\) are input features, the components of an input vector \(\mathbf{x}\) describing a data point. The weights \(w_1, w_2, w_3\), collected into a weight vector \(\mathbf{w}\), are learned parameters that determine how much each feature contributes to the output. The weighted sum and bias are now written compactly as the pre-activation \(\mathbf{w}^\top \mathbf{x} + b\), and applying the activation function \(f\) yields the output \(y = f(\mathbf{w}^\top \mathbf{x} + b)\). When \(f\) is the identity function, this is exactly the linear model we studied in earlier chapters.

Both perspectives describe the same computation, which we can write compactly as:

\[y = f(\mathbf{w}^\top \mathbf{x} + b)\]

  • \(\mathbf{x}\): input vector or features
  • \(\mathbf{w}\): weight vector
  • \(b\): bias
  • \(f\): activation function
  • \(y\): the neuron’s output (or activation)

Question: In the simplified neuron, the cell body accumulates the total signal \(\sum_j w_j x_j + b\) and passes it through the activation function \(f\). What does \(f\) correspond to in a biological neuron?

Answer:

The activation function captures the idea that a biological neuron’s output is not simply the sum of its incoming signals. Depending on the total signal it receives, the neuron may fire or stay quiet. In the model, \(f\) plays this role: it takes the accumulated signal \(\sum_j w_j x_j + b\) and determines whether and how strongly the neuron fires. We will see in a later chapter that choosing a non-linear \(f\) is essential to the expressiveness of neural networks.

Question: A biological neuron’s output is often thought of as binary: the neuron either fires or it does not. Is the output of an artificial neuron also binary?

Answer:

Not usually. In most cases, we choose the activation function \(f\) so that the output of an artificial neuron is a real number rather than just \(0\) or \(1\). We can think of this real-valued output as emulating the firing rate of a biological neuron (how rapidly it fires) rather than the all-or-nothing decision of whether it fires at a single instant.

4 Logistic Regression as a Neuron

This simplified model (weighted inputs, a bias, and an activation function) is the basic building block of neural networks. It should look familiar. In fact, if the activation function \(f\) is the sigmoid function \(\sigma\), this neuron is exactly logistic regression. So you can think of logistic regression as a single neuron.

When denoting logistic regression as a neuron in a neural network, we often like to represent it using a graph. Figure 3 represents logistic regression as a directed graph and shows how information flows through the model. Figure 4 shows the same model as a mathematical formula.

Logistic regression neuron

Figure 3

Neuron formula

Figure 4: Interactive Hover (or tab to focus) each symbol (\(y\), \(f\), \(\mathbf{w}^\top\), \(\mathbf{x}\), or \(b\)) to see what it represents.

Consider the diagram of a single neuron in Figure 3. Suppose it has three inputs, \(x_1, x_2, x_3\). Each circle represents a value, just a number. These inputs can be thought of as signals coming from other neurons. The arrows represent connections, and each connection is labelled with a weight that indicates the strength of that connection. Finally, the neuron produces an output, often denoted by \(y\).

What is not shown explicitly in Figure 3 is the computation itself. The diagram shows the inputs, the weights, and the output, but it does not explicitly depict the steps of multiplying each input by its weight, summing the results, adding the bias, and applying the activation function. All of that computation is implied. The diagram tells us what is connected to what, not how each numerical step is carried out.

Figure 3 is an example of a network diagram. Compare it with Figure 2: both show the same neuron computing \(y = f(\mathbf{w}^\top \mathbf{x} + b)\), but at different levels of detail. Figure 2 draws the inside of the neuron, including the synapses carrying the weights, the weighted sum, and the activation function. Figure 3 collapses all of this into a single node, and each synaptic connection becomes an arrow.

The network diagram also shows something that Figure 2 leaves out. In Figure 2, the inputs \(x_1, x_2, x_3\) are just signals arriving along incoming axons. In Figure 3, each input gets its own node, as if it were the output of an upstream neuron or a sensor. This is a drawing convention, not a change to the model. We explore this and other conventions, such as drawing the bias as a node, in Network Diagram Conventions.

Network diagrams show the overall structure of a model. Later, we will introduce computation graphs, which make each numerical step explicit and are useful for reasoning about gradients. We discuss them in Backpropagation.

5 Softmax Regression as a One-Layer Neural Network

Now that we know how to represent a single neuron, we can connect this back to even more models you already know. Recall that softmax regression is a generalization of logistic regression to the multi-class setting. In logistic regression, there are two classes, so we only need to model the probability of one of them. In softmax regression, we have \(K\) classes, and we want to produce one probability for each class.

Softmax one-layer network

Figure 5: The model is \(\mathbf{y} = \text{softmax}(\mathbf{W}\mathbf{x} + \mathbf{b})\), where \(\mathbf{W}\) is the weight matrix and \(\mathbf{b}\) is the bias vector.

Softmax regression can be viewed as a one-layer neural network for multi-class classification. It consists of multiple output units, one per class, and uses the softmax function to convert raw scores into a probability distribution over the classes.

Let’s look at the softmax regression structure with this new perspective in mind. To compute the score for each class, the model multiplies each input by its corresponding weight through a matrix-vector product and adds a bias vector. It then applies the softmax function to convert the class scores to a probability distribution over the classes. As before, these computations are implied but not shown explicitly in the network diagram.

Therefore, linear models, such as logistic regression and softmax regression, can all be interpreted as neural networks with a single layer. This perspective will serve as a foundation as we move on to deeper and more expressive models.

6 Network Diagram Conventions

Different conventions are sometimes used to represent the same underlying neural network. The diagrams in this chapter have already varied in a few independent ways. For example, Figure 2 and Figure 3 represent the same neuron with the same underlying mathematical structure.

In the rest of this section, we will use several conventions for drawing neural network diagrams. We sometimes use a vertical layout, with inputs at the bottom, outputs at the top, and arrows pointing upward. Other times, we use a horizontal layout, with inputs on the left, outputs on the right, and arrows pointing from left to right. These two views are equivalent and differ only in presentation.

Likewise, the bias is sometimes drawn as an explicit node, so that the bias becomes a weight annotated on an arrow. This is analogous to the linear regression section, where we treated the bias parameter as an extra weight. Figure 6 lets you pick each of these independently and see the resulting diagram. Given the same number of output units, all other choices do not change the underlying network being represented, only the representation.

Network diagram conventions

Figure 6: Interactive Choose the representation settings of a one-layer neural network with 1, 2, 3, or 4 output units. The remaining settings (whether to show the inputs as nodes, the bias as a node, whether to annotate the weights, and whether to orient the layers horizontally or vertically) do not change the underlying network being represented. An arrow in this diagram represents a synaptic connection; mouse over an arrow to highlight the corresponding weight (if weight annotation is shown).

When a layer has many units, drawing every node and arrow becomes cluttered. A more compact convention draws each layer as a single rectangle labelled with its size or operation, with one arrow between consecutive layers standing for all of their connections. This style is common in research papers.

7 Summary

An artificial neuron abstracts the basic mechanisms of a biological cell: dendrites receive inputs, the cell body integrates them, and the axon carries an output to downstream neurons. Synaptic strengths become learnable weights, and the firing decision becomes an activation function. The resulting unit computes \(y = f(\mathbf{w}^\top \mathbf{x} + b)\).

Viewed this way, the linear classifiers from earlier chapters are already neurons or single-layer neural networks. This makes clear both what these models compute and why they are limited, namely the same constraint analyzed in Limitations of Linear Models: a single affine map followed by an activation can produce only linear decision boundaries. The next step is to connect many neurons into deeper networks, starting with multi-layer perceptrons, and to study how depth and nonlinearity unlock expressiveness that single-layer models lack.