Mathematics & Logic

Thinking in lists of numbers

Vectors and matrices, the school maths that search engines, recommendations and every neural network quietly run on.

  • 8min read
  • 10min listen
  • 29questions
Five slender paper arrows fan out from one point; two blue arrows point almost the same way, side by side.

Thinking in lists of numbers

0:00 / 10:11

A question to hold while you read

A computer only handles numbers. So how can it judge that two sentences mean nearly the same thing?

Everything as a list

A computer cannot look at a house the way a buyer does. It can only work with numbers. So machine learning describes each thing as an ordered list of numbers. A house might become (3, 120, 25): three bedrooms, 120 square metres of floor and 25 years of age. Such a list is called a vector.

The order matters, because each position means one feature, always the same one. Swap the first and last numbers and you get (25, 120, 3), a very different house with 25 bedrooms that is only 3 years old. Once a house, a photo or a word has been turned into a vector, the same few operations work on all of them: adding, stretching and, above all, comparing. That is why this branch of school maths, called linear algebra, became the everyday language of artificial intelligence.

A point, or an arrow

A list of two numbers can be drawn. Take (3, 4): start at the corner of a grid, go 3 steps across and 4 steps up, and mark the spot. The vector is that point, or equally the arrow from the corner to it. An arrow has a direction as well as a size, which is why physics draws forces and velocities as vectors too.

The arrow's length comes from Pythagoras. The arrow is the long side of a right-angled triangle whose other two sides are 3 and 4, so its length is the square root of 3² + 4², the square root of 25, which is 5. The same recipe works for a list of any size: square every entry, add the squares, then take the square root. The distance between two points comes out the same way, from the differences between their entries.

93 across, squared164 up, squared25length, squared
Pythagoras for the arrow (3, 4): 9 + 16 = 25, and the square root of 25 is 5

Beyond three dimensions

Why stop at two or three numbers? A small greyscale photo, 28 pixels wide and 28 tall, can be stored as 784 numbers, one for the brightness of each pixel, read off row by row. That list is a vector, and the photo is a single point in a space with 784 dimensions, one for each entry.

Nobody can picture 784 directions, each at right angles to all the others, and nobody needs to. The arithmetic does not care. Length, distance and every other operation in this Topic work in exactly the same way with 784 entries as with two. Machine learning lives in spaces like this. Google's widely shared word2vec lists give each word 300 numbers, and a large language model uses thousands of numbers for each piece of text it reads.

3 numbersa house300 numbersa word in word2vec784 numbersa small photo
How many numbers describe one of each, in this passage's examples

Adding and stretching

Vectors add entry by entry: (3, 1) plus (1, 2) is (4, 3). Drawn as arrows, that is a walk. Go 3 across and 1 up, then from where you stand go 1 across and 2 up. With the second arrow's tail set on the first arrow's tip, the sum is the single arrow from start to finish.

Multiplying by an ordinary number stretches a vector instead. 2 × (3, 1) is (6, 2): the same direction, twice as long. A half shrinks it, and a negative number flips it to point the opposite way. A plain number used like this, with a size but no direction of its own, is called a scalar, because it scales. The two moves combine, too. The point halfway between two points is their sum scaled by one half, which is how an average of many vectors is found.

acrossup
Two short sides: walk (3, 1), then (1, 2). The long side, from start to finish, is the sum (4, 3)

3 more questions from this passage

Multiply and add: the dot product

One operation does most of the work in machine learning. Take two lists of the same length, multiply their matching entries, and add up the results. The single number you get is the dot product of the two lists, named after the dot written between them.

A shopping trip is one. You buy 2 loaves and 3 litres of milk, so your quantities are (2, 3). Bread costs 1.5 a loaf and milk 1 a litre, so the prices are (1.5, 1). The dot product is 2 × 1.5 + 3 × 1, which is 6: your bill. An artificial neuron does the same sum. It multiplies each input by a weight, a number saying how much that input counts, then adds the results. That weighted sum is simply the dot product of the inputs with the weights.

1 more question from this passage

A score for pointing the same way

Drawn as arrows, the dot product measures agreement. Two arrows pointing the same way give a large positive number: (1, 1) with (2, 2) gives 1 × 2 + 1 × 2, which is 4. Turn one arrow and the score shrinks as the angle between them opens. At exactly a right angle, as with (1, 0) and (0, 1), every product is zero, so the dot product is zero, and the arrows are called perpendicular. Past a right angle the score turns negative, and arrows pointing in opposite directions, like (1, 1) and (−1, −1), give the most negative score their lengths allow.

So one multiply-and-add answers a useful question about any two lists of numbers. Do they rise and fall together, have nothing to do with each other, or pull against each other?

2 more questions from this passage

Similar meaning, measured

The raw dot product has a flaw: it also grows with length. A long arrow scores high against almost everything, just for being long. The fix is to divide the dot product by the two arrows' lengths, so that only the angle between them counts. What is left is the cosine of that angle, called cosine similarity: 1 for the same direction, 0 at right angles, −1 for opposite directions. The arrows (1, 1) and (2, 2) score exactly 1, though one is twice as long.

This is how software judges similar meaning. A search engine or a chatbot turns each passage of text into a long list of numbers, an embedding, built so that texts about the same thing point in similar directions. Your question gets one too, and the passages with the highest cosine similarity to it come back first.

-1-0.500.51oppositeright anglessame direction
The range of cosine similarity, from minus one to one

That answers the question you started with: A computer only handles numbers. So how can it judge that two sentences mean nearly the same thing?

3 more questions from this passage

Taste as a list of numbers

In 2006 Netflix offered a million dollars to anyone who could predict viewers' film ratings 10 per cent better than its own system. The team that won in 2009 leaned heavily, among other methods, on dot products. Give every film a short list of numbers, and every viewer a list of the same length. One entry might score how serious rather than escapist a film is, and the viewer's matching entry how much they enjoy serious films. The viewer's predicted rating is, at heart, the dot product of the two lists.

Nobody types these numbers in. The system starts from guesses and adjusts them until the dot products match the millions of ratings people have already given, so the latent factors, hidden tastes like serious against escapist, are learned from past ratings. Researchers can sometimes read a meaning into one afterwards, but nobody chooses them in advance.

2 more questions from this passage

King minus man plus woman

Word vectors seemed to hold meaning in their directions. In 2013 Tomas Mikolov and colleagues reported that king − man + woman, worked out entry by entry, lands nearest to queen, and that Madrid − Spain + France lands near Paris. The word analogy became the most famous demonstration in the field.

The trick is weaker than the story. The search quietly leaves out the three words in the question. Allowed to answer with them, the nearest word to king − man + woman is king itself: the arithmetic barely moves the point. In 2020 three researchers at the University of Groningen measured this on Mikolov's own analogy test. Correct answers fell from 74 per cent to 21 once the question's words were allowed back. Grammar, like singular to plural, also works far better than meaning, like opposites. The famous examples are the method at its best.

74 per centquestion words excluded21 per centquestion words allowed
Analogy questions answered correctly, Google News word vectors (Nissim, van Noord and van der Goot, 2020)

2 more questions from this passage

A grid of numbers

Stack many vectors together and you get a matrix: a rectangle of numbers in rows and columns. A spreadsheet of 100 houses, each a row of 3 numbers, is a matrix with 100 rows and 3 columns.

The idea is old. The Nine Chapters on the Mathematical Art, compiled in China roughly two thousand years ago, sets out problems with several unknowns, such as the yields of three grades of grain, as rectangles of counting rods on a board. It then clears the unknowns one at a time by combining whole columns, the method Europe much later named after Carl Friedrich Gauss. The English name came in 1850, when James Joseph Sylvester called such an array a matrix. In 1858 Arthur Cayley went further and treated matrices as things in their own right, to be added and multiplied.

2 more questions from this passage

A machine that moves space

A matrix can also act on a vector. To multiply a matrix by a vector, take the dot product of each row with the vector, and each answer becomes one entry of a new vector. The matrix with rows (0, −1) and (1, 0) sends (3, 1) to (−1, 3), the same arrow given a quarter turn anticlockwise. It turns every other arrow the same way, so it is a rotation. The matrix with rows (2, 0) and (0, 1) stretches everything to twice its width.

So a matrix is a machine for a linear transformation: it rotates, stretches, flips or squashes the whole space at once while keeping straight lines straight. Doing one transformation and then another gives the same result as multiplying the two matrices together, the rule Cayley set out in 1858, so a long chain of moves collapses into a single matrix.

3 more questions from this passage

A layer of a neural network

Here the pieces meet. Feed a network the 784 numbers of a photo as one vector. Its first layer might hold 100 artificial neurons, each with 784 weights, one per pixel. Stack those weights as rows and you have a matrix of 100 rows and 784 columns. One matrix multiplication then gives all 100 weighted sums at once, each the dot product of one neuron's row with the photo.

Then comes a simple bend, called the activation function. A common one turns every negative number into zero and leaves the rest alone. The bend is essential. Without it, a stack of layers is only a chain of matrix multiplications, and a chain of matrices collapses into one matrix. Fifty layers without bends would do the same as one layer. With a bend after each, every layer can build on the last.

3 more questions from this passage

Why graphics chips took over

A video game redraws its world dozens of times a second. Every corner of every shape on screen is a vector, and to show the scene from the player's viewpoint each one is multiplied by the same small matrices, for every frame. Chips built for that job, graphics processing units or GPUs, have thousands of simple cores doing multiply-and-add side by side. Nvidia sold its GeForce 256 as the first GPU in 1999.

A neural network layer is the same kind of job, a matrix multiplication whose outputs are separate dot products, none waiting for another. Both come down to matrix multiplication. In 2009 Rajat Raina, Anand Madhavan and Andrew Ng trained large networks on a graphics card up to 70 times faster than on an ordinary two-core processor, cutting weeks of training to about a day. Chips made for games became the engine of AI.

1two-core processor70GPU
Training a large network in 2009: best-case speed relative to an ordinary two-core processor

2 more questions from this passage

29 questions came out of this reading. Answer them out loud on your phone, and EdenMind schedules each one for the day you’re about to forget it.

Add to my practice