Unsupervised Learning

1 Introduction

For all of the methods that we have examined so far, we have been interested in predicting the value of one target variable from the values of the other variables. For classification problems, the target has been a categorical class label (e.g. oak, birch, or maple) and for regression problems, the target has been a continuous variable (e.g. the temperature in Celsius of the UTM pond). In every case, we have had a labelled training dataset consisting of points with known target values.

But what if we don’t have a labelled training set? What if we either don’t have the target values for our data, or there is no particular target variable that we want to predict? Can we still do useful machine learning?

It turns out that the answer is yes. We can perform unsupervised learning. In contrast with the supervised learning we have previously done that has required labelled training data, unsupervised learning works on unlabelled data—data for which there are no target values.

Since there are no labels, unsupervised learning is used to find structure in the data in a way that is fundamentally different from the predictors we get through supervised learning. In this class we will look at two useful types of unsupervised learning:

  • Clustering: Do data form clusters? (I.e. Are there groups of data points that are more similar to each other than to other data points?) This is the subject of the current chapter.

  • Dimensionality Reduction: Can data be expressed in a lower-dimensional space? (I.e. Do data points tend to fall in a subspace or manifold of the space that they inhabit?) Principal Component Analysis (PCA), which can be used for dimensionality reduction, is the subject of the next chapter.

2 Clustering

In clustering, we want to identify collections of points that are more similar to each other than they are to other points, where similarity can be taken to mean closeness in the data space. Although it will be convenient to work in two-dimensional space in this section for ease of visualization, the clustering methods we discuss may be applied to data in any number of dimensions.

Consider the following collection of points, which represent leaves whose species we do not know:

Unlabelled Leaf Data

Figure 1: Sixty unlabelled leaves are plotted by width and height.

Since we do not have any notion of class label for these points, all are plotted using markers of the same shape and colour. Even so, perhaps you can intuitively observe some structure in this data. If you were to group these data into clusters based on their closeness in this two-dimensional space, how many clusters would you make? What would be the (x,y) coordinates of the center points of those clusters?

Below, we have given two possible clusterings of this data, first for two clusters, and then for three. The markers of points within the same cluster share a shape and colour.

Outlined in bold black, we have also drawn the centroid of each cluster. The centroid is not an actual data point in the dataset. Rather, it gives the average coordinates of all points within the cluster (i.e. the center of mass).

Unlabelled Leaf Data - 2 Clusters

Figure 2: The same leaf dataset with two clusters identified. Each point is coloured and shaped according to its assigned cluster; centroids are shown with a bold black outline.

Unlabelled Leaf Data - 3 Clusters

Figure 3: The leaf dataset with three clusters identified. Each point is coloured and shaped according to its assigned cluster; centroids are shown with a bold black outline.

Both of these clusterings have been optimized (using a Gaussian Mixture Model, to be discussed later in this chapter), but they have been optimized for different numbers of clusters. We might hope that these different numbers of clusters can help us to interpret the data at different levels of granularity (e.g. individual species of trees versus family of trees), but in general, there is no guarantee that the clusters are meaningful. Perhaps the leaves were simply collected in two or three different seasons, or perhaps there is no human-friendly explanation for the clusters.

Ultimately, what we can infer from a cluster is that the points within tend to be more similar to each other than they are to the points outside the cluster, and that can be enough to be useful. For example, we might test a sample of leaves in each cluster for disease, giving us an estimate of the occurrence of disease within the cluster as a whole. Then, when we find a new leaf (one that is not part of the original dataset), we might check to which cluster it ought to belong, and use that information to decide whether or not to test for disease.

3 Summary

We have introduced the notion of unsupervised learning, which, unlike the supervised learning we have done so far, does not require labelled training data, and also the notion of clustering, a type of unsupervised learning that finds clusters of similar data points.

For the remainder of this chapter, we will examine two related methods for producing clusterings: K-means and Gaussian Mixture Models (GMMs).