Mathematics & Logic

Learning with no one to mark the answers

Labels are slow and costly to make. So machines learned to set their own questions: group what is alike, rebuild what was squeezed, guess what was hidden.

  • 8min read
  • 11min listen
  • 26questions
Small paper pebbles on off-white paper, gathered into three loose clumps, one clump violet.

Learning with no one to mark the answers

0:00 / 10:31

A question to hold while you read

Nobody has labelled the billions of sentences on the internet. So how can a machine learn anything from them?

The price of a label

Most machine learning learns from labelled examples: an X-ray marked "fracture", a product review marked as praise or complaint. Someone has to supply every one of those answers. ImageNet, the photo collection behind a turning point in computer vision, holds more than 14 million labelled images. Sorting them took about 49,000 paid online workers in 167 countries, hired through Amazon's Mechanical Turk service over about two years, starting in 2008.

Labels on that scale are rare and slow to make. Unlabelled data is everywhere: photos, recordings, shop receipts, and whole libraries of text that nobody has marked. Learning from data with no answers attached is called unsupervised learning. With nothing to check its answers against, the program has to find structure in the data itself: groups, summaries, or parts that predict one another.

Groups nobody named

Picture a supermarket's till records: millions of baskets, and nobody has said what kinds of shopper there are. Place each customer by what they buy, and some sit close together. Here are baskets heavy on nappies and baby food; over there, instant noodles and energy drinks. A program that finds such groups on its own is doing clustering. A cluster is a set of points that lie nearer to one another than to the points outside it.

Shops use this for customer segmentation: splitting shoppers into groups and aiming a different offer at each. The groups are found, not given. Nobody wrote "young parents" into the data. The program found a crowd of similar baskets, and a person looked at it and gave it that name.

The k-means recipe

The best-known way to find clusters is called k-means, where k stands for the number of groups you ask for. Its recipe is short. Drop k centre points among the data, anywhere at all. Then repeat two steps. First, every data point joins its nearest centre. Second, every centre moves to the average position of the points that joined it. Moved centres pull in slightly different points, so the two steps run again, and again, until no point changes group.

Stuart Lloyd of Bell Labs worked out this procedure in 1957, to choose the few levels a telephone signal could be rounded to when it was sent as numbers. His note was not published until 1982, and the statistician James MacQueen gave the method its name in 1967. Each round can only tighten the groups, so the recipe always settles, though not always on the best grouping possible.

points join nearest centrecentres move to average
The two steps k-means repeats until no point changes group

How many groups?

k-means never asks whether the data holds any groups. Give it pure noise and a k of five, and it returns five neat clusters. The number of groups is a choice the person makes before the recipe runs.

Could you let the data choose, by picking the k that gives the tightest groups? No, because the groups always get tighter as more clusters are added. Give every point a cluster of its own and the spread inside each is zero, with nothing learned. So people plot the spread against the number of clusters and look for the elbow: the bend where the curve stops falling steeply, and extra clusters stop buying much. Often the bend is soft, and two people read it differently. The starting centres matter too. Different random starts can settle on different groupings, so the recipe is usually run several times and the tightest result kept.

clustersspread
Spread inside the groups as more clusters are asked for

2 more questions from this passage

Squeeze, then rebuild

Another way to learn without labels is to make a network copy its input. An autoencoder has two halves. The encoder squeezes each input down through a narrow middle layer, the bottleneck. The decoder rebuilds the original from what came through. Training nudges the network until copy and input match. The input is its own answer, so nobody has to label anything.

Copying sounds pointless, but the bottleneck makes it hard work. In 2006 Geoffrey Hinton and Ruslan Salakhutdinov trained one on handwritten digits, 784 pixels each, squeezed through a middle of just 30 numbers. To rebuild a digit from 30 numbers, the network had to capture the main ways handwritten digits differ, the sort of thing a person might call slant, loops or width. Those 30 numbers became a compact summary of each image, learned from the pictures alone, and they rebuilt digits better than an older standard method.

784 numbersimage in30 numbersbottleneck784 numbersrebuilt copy
Numbers describing one handwritten digit, in and out

3 more questions from this passage

Known by the company it keeps

Words pose the same puzzle. How could a program learn what a word means when nobody tells it? In 1957 the British linguist J. R. Firth put the answer in one line: "You shall know a word by the company it keeps."

Try it. You may never have met the word sahlab. But read "she stirred the hot sahlab and poured it into two cups" and "a mug of sahlab on a cold evening", and you can tell it belongs with tea and cocoa, not with bricks. The words around it gave it away. Words that turn up in similar surroundings tend to have similar meanings. This is the distributional hypothesis, and it turns meaning into something a machine can count: which words appear near which others, across millions of sentences.

1 more question from this passage

Learning from the neighbours

In 2013 a Google team led by Tomas Mikolov turned Firth's idea into a fast training method called word2vec. Its best-known version, skip-gram, slides along ordinary text. At each word it tries to predict the words around it, a few places before and after. Every word starts as a list of random numbers, and every wrong guess nudges the lists a little.

Nobody labels anything: the text supplies both the question and the answer. After billions of words, "tea" and "coffee" end up with similar lists, because they have similar neighbours, and the only way to predict those neighbours well is to treat the two words alike. The lists are the real prize. They are embeddings, and words used in similar ways end up close together among them. The guessing game was only a way to force the numbers to capture how each word is used.

2 more questions from this passage

Make the data mark itself

Word2vec points to a bigger trick. Take any data, hide part of it, and train a model to guess what was hidden. The hidden part is the answer, so the data supplies its own labels. This is self-supervised learning, a branch of unsupervised learning, and it is how the internet became training material.

In 2018 Google's BERT applied it to sentences. It hid about 15 per cent of the words and learned to fill each gap from the words on both sides, reading 3.3 billion words of books and Wikipedia. The models that write chat replies use a close cousin of the task: guess the next word from everything before it. Neither game needs a human marker. Every sentence ever written becomes practice, and a model that fills gaps well has to pick up a good deal about grammar, facts and how words relate.

That answers the question you started with: Nobody has labelled the billions of sentences on the internet. So how can a machine learn anything from them?

2 more questions from this passage

Hiding most of a picture

Pictures can play the same game. In 2021 Kaiming He and colleagues at Facebook AI Research cut images into a grid of small square patches, hid a random 75 per cent of them, and trained a masked autoencoder to paint the missing pixels back in from the quarter that remained.

Why hide so much, when BERT hid roughly one word in seven? Neighbouring pixels are nearly alike, so with only a few patches gone a model could fill each gap by smudging in the colours beside it, learning nothing about what the picture shows. Hide three quarters and smudging fails. To rebuild a dog's missing head from a paw and a tail, the model has to learn what dogs look like. Pictures repeat themselves far more than sentences do, so the puzzle has to be made harder.

020406080100words, BERTimage patches
Share of the input hidden during training, in per cent

2 more questions from this passage

Pull together, push apart

Rebuilding every pixel is not the only way to learn from pictures without labels. Contrastive learning asks a simpler question: which of these belong together?

In 2020 Ting Chen, Geoffrey Hinton and colleagues at Google showed a version called SimCLR. Take a photo and make two altered versions of it, cropped differently and with the colours shifted. These two copies of one photo are a matching pair, and every other photo in the batch is a mismatch. The network turns each image into a list of numbers, and training pulls the lists of a matching pair together while pushing them away from all the others. To succeed, the network has to ignore what cropping and colour shifts change, and keep what they leave alone, such as the shape of a cat's ear. Nobody labelled a single photo, yet the lists it learned captured what the pictures show.

2 more questions from this passage

Captions the web already wrote

The internet also pairs pictures with words. In 2021 OpenAI trained a system called CLIP on 400 million images collected from the web, each with the text that came with it. Its task was contrastive: given a batch of images and a batch of captions, work out which caption matches which image.

People wrote those captions for their own reasons, never as labels for a training set, so nobody was paid to mark anything. The payoff was flexibility. To sort photos into ImageNet's 1,000 categories, CLIP simply matched each photo against phrases naming each category, such as "a photo of a dog". Without training on any of the 1.28 million labelled images made for that contest, it matched the accuracy of ResNet-50, a standard network trained on all of them.

1.28 millionImageNet labelled400 millionCLIP image-text pairs
Training examples in millions, labelled by hand or gathered from the web

2 more questions from this passage

What babies pick up

People find structure without labels too, and one experiment showed how early. In 1996 Jenny Saffran, Richard Aslin and Elissa Newport played 8-month-old babies two minutes of a made-up language: four three-syllable words such as tupiro and golabu, strung together in a flat synthetic voice with no pauses. Within words, each syllable was always followed by the same next one. Between words, any of three syllables could follow, so each pairing held only one time in three.

Afterwards the babies listened longer to strings that were not words of the language, showing that they could tell the words apart. Picking up which sounds tend to follow which is called statistical learning. It is one ingredient of learning a language, not the whole of it: babies also have voices they love, faces to watch, and a world the words are about.

100 per centwithin words33 per centbetween words
How predictable the next syllable was, in per cent

2 more questions from this passage

Patterns, not meanings

Learning without labels removes the costly human step, and with it a check. Whatever patterns fill the data, the model absorbs. In 2017 Aylin Caliskan, Joanna Bryson and Arvind Narayanan examined word embeddings learned from 840 billion words of web text. The associations psychologists measure in people were there. Female names sat closer to family words and male names to career words, and flowers sat closer to pleasant words than insects did. Nobody put those stereotypes in on purpose. They came with the text, and the biases in the text became biases in the model.

A pattern is also not a meaning. A cluster of baskets is not yet a kind of shopper, and a word's neighbours are not the thing the word names. These methods find what goes with what, often remarkably well. Deciding what the patterns mean, and whether they are fair, still falls to people.

2 more questions from this passage

26 questions came out of this reading. Answer them out loud on your phone, and EdenMind schedules each one for the day you’re about to forget it.

Add to my practice