Mathematics & Logic

Changing your mind by the numbers

Start from what you already knew, weigh the new clue, and end up exactly as sure as the evidence allows.

  • 9min read
  • 10min listen
  • 28questions
A paper counting frame with cream beads on five rods, three violet beads slid apart on the top rod.

Changing your mind by the numbers

0:00 / 9:52

A question to hold while you read

A test that catches most cases of a rare cancer comes back positive. Why is the woman who took it still probably healthy?

Where a belief starts

A patient walks into a clinic with a cough. Before the doctor has listened to the chest, which is more likely: a cold or pneumonia? A cold, by a long way, simply because colds are far more common. A good doctor carries that knowledge into the room before any test is run.

That starting judgement, made before the new evidence arrives, is called the prior. Often it comes from a base rate: how common each explanation is among people like this one. The prior is not a prejudice to be thrown away. It sums up everything learned so far, and every sensible conclusion begins from one. The question this Topic answers is how far a new piece of evidence should move it.

How much a clue is worth

Now the doctor listens to the lungs and hears a crackling sound. Should this change the verdict? Ask a sharper question: how probable is this finding if the patient has pneumonia, and how probable if it is a cold? Crackling lungs are common in pneumonia and rare with a cold, so they count heavily.

How probable a piece of evidence is under a given explanation is called its likelihood. What matters is the comparison between explanations. The cough itself is likely under both, so it hardly separates them and should barely move the doctor's belief. A clue is strong evidence only when it is much more likely under one explanation than under the other. The bigger that gap, the further the belief should shift.

Belief after the evidence

Put the two pieces together and you have the whole method. Begin with the prior, weigh each explanation by how likely it makes the evidence, and compare. The result is the posterior, your belief after the evidence. A rare explanation can win if the evidence favours it strongly enough, and a common one can win despite some evidence against it. This recipe is called Bayes' rule.

It also runs in steps. Today's posterior becomes tomorrow's prior. The doctor who now suspects pneumonia orders an X-ray, and the X-ray updates that belief again. Nothing has to be settled once and for all. Each new piece of evidence moves the belief a little or a lot, depending on its strength, and the next conclusion starts from wherever the last one left off.

priorevidenceposterior
Bayes' rule, run again with each new clue

A minister's unpublished essay

The rule carries the name of Thomas Bayes, an English Presbyterian minister with a gift for mathematics, who never published it. After he died in 1761, his friend Richard Price, another minister, found an essay among his papers, edited it and added to it, and read it to the Royal Society in London in December 1763.

Bayes had tackled a backward question. The mathematics of chance already ran forwards: given a fair coin, how often will it land heads? Bayes asked the reverse. Given only how many times an event has happened and failed, what can we say about the chance behind it? Reasoning back from what happened to the chance that produced it came to be called inverse probability. It is the problem every learner faces, because the world shows us only outcomes; the odds behind them stay hidden.

1760178018001820Price reads Bayes's essayLaplace's own versionthe sunrise example
How Bayes's idea reached the world

Will the sun rise tomorrow?

Bayes's idea was built into a general method by the French mathematician Pierre-Simon Laplace, who worked it out for himself in 1774, apparently without knowing of Bayes's essay. In 1814 he gave a famous example. Suppose you know no astronomy, only that the sun has risen on every one of 1,826,213 days, about five thousand years of recorded history. How sure should you be about tomorrow?

Laplace's answer, for a starting point of total ignorance, now called the rule of succession, is to add one to the successes and two to the tries. If something has happened every time in three tries, the chance it happens next is four in five, or 80 per cent. For the sun, the odds come to 1,826,214 to 1. But Laplace added that someone who understands the cause, what drives the days and seasons, has far better grounds than the count alone.

50 %no tries yet67 %1 of 180 %3 of 392 %10 of 1099 %100 of 100
Laplace's rule: the chance of one more success, when every try so far has succeeded

2 more questions from this passage

A positive test, counted out

Here is the example that trips up most people, worked with counts instead of percentages. The numbers are the textbook ones used in a well-known 1995 study, not today's screening statistics. Picture 1,000 women aged forty at a routine breast screening. Ten of them have breast cancer, and of those ten, 8 test positive. Of the 990 women without cancer, 95 also test positive.

Now take one woman whose result came back positive. How worried should she be? Count everyone with a positive result: 8 plus 95 makes 103. Only 8 of those 103 actually have cancer, fewer than one in twelve. The test is useful, since a positive result raises her chance from 1 in 100 to about 8 in 100. But most positive results are false positives: false alarms among the many healthy women.

That answers the question you started with: A test that catches most cases of a rare cancer comes back positive. Why is the woman who took it still probably healthy?

2 more questions from this passage

Two questions that sound alike

Why does that answer feel so wrong? Because two different questions hide behind similar words. The test finds 8 of every 10 cancers: that is the chance that a woman with cancer tests positive. The woman holding a positive result wants the reverse: the chance that she has cancer, given that she tested positive. The first is 80 per cent; the second is under 8.

What separates them is the base rate, how common the condition is to begin with. Cancer was present in 1 woman in 100, so the healthy outnumber the sick by 99 to 1. Even a small rate of false alarms, applied to that crowd of healthy women, produces more positive results than the few real cases can. Ignoring the base rate is one of the commonest errors in reasoning about evidence, made by students and doctors alike.

2 more questions from this passage

Counts, not percentages

The same problem is usually posed in percentages: the cancer affects 1 per cent of women, the test detects 80 per cent of cancers, and it wrongly flags 9.6 per cent of healthy women. In one often-cited 1982 report, most doctors asked put the chance that a woman with a positive result has cancer between 70 and 80 per cent.

In 1995 Gerd Gigerenzer and Ulrich Hoffrage argued that much of the trouble lay in the format. They gave 60 students at the University of Salzburg fifteen problems of this kind. When the facts came as percentages, 16 per cent of the answers followed the correct reasoning. Given the same facts as counts of people, such as 10 out of 1,000, the share rose to 46 per cent. Such counts came to be called natural frequencies: the numbers you would meet by watching cases one at a time.

16 %percentages46 %natural frequencies
Answers that reasoned correctly, by how the same facts were stated (1995)

2 more questions from this passage

A filter that learns from words

In August 2002 the programmer Paul Graham published an essay called "A Plan for Spam". Instead of writing rules about what junk mail looks like, he let the mail itself supply the evidence. He kept two collections, about 4,000 spam messages and 4,000 wanted ones, and counted how often each word appeared in each.

A word like "madam" or "promotion" turned up almost only in spam, so it pushed a new message toward spam. Words like "though" and "tonight" pushed the other way: each was common in his real mail and rare in spam. The filter combined the evidence of a message's most telling words with Bayes' rule. Filters of this kind treat each word as a separate clue, ignoring how words go together, and so are called naive Bayes. On his own mail, Graham reported missing fewer than 5 spams in 1,000.

2 more questions from this passage

A probability, then a decision

Graham's filter did not say "spam" or "not spam". It said how probable spam was, and he explained why that mattered. Rival filters gave each message a score, and nobody, he wrote, knew what a score meant. A probability has a meaning you can check.

The verdict comes afterwards, from a threshold: a cut-off above which a message is treated as spam. Graham set his at 90 per cent. Where to put it is a question of cost, not evidence. For most people, he argued, losing a real email is worse than seeing some junk, roughly ten times worse, so the filter should be very sure before it hides anything. A medical test, a fraud alarm or a network sorting pictures works the same way: first a probability, then a decision that weighs what each mistake would cost.

2 more questions from this passage

When 70 per cent means 70

If a model says "70 per cent", how would you know whether to believe it? Not from one case: a single message is either spam or it is not. Instead, collect every case where the model said 70 per cent and count how many of them were right. A trustworthy model is right about 70 times in 100 on those cases, and about 90 in 100 when it says 90. That match between stated confidence and how often it comes true is called calibration.

Weather forecasts are judged this way. Suppose a forecaster says "90 per cent chance of rain" on many days, but it rains on only about 60 of every 100 such days. Her forecasts are poorly calibrated, even if she is right on most days. Calibration is separate from accuracy.

1010%3030%5050%7070%9090%
Days it rained, out of 100, for each forecast chance (a calibrated forecaster)

2 more questions from this passage

Sure of themselves

In 2017 Chuan Guo and three colleagues at Cornell University checked the calibration of networks that recognise pictures. They compared a small network from 1998 with a 110-layer network from 2016, on a test with 100 kinds of picture. The modern network was more accurate: it got about 31 per cent wrong, against 45 per cent for the old one. Yet the old network was well calibrated and the new one was not. Its confidence ran well above its accuracy. It was overconfident.

Their recommended fix was simple. Before a network turns its raw scores into probabilities, divide every score by one number, called the temperature, tuned on examples held back from training. This softens the probabilities without changing which answer the network picks. The method is called temperature scaling.

2 more questions from this passage

Every learner starts somewhere

A learning machine needs a prior too. For any word his filter had never seen, Graham set a spam probability of 0.4, a small lean toward innocence, because, as he put it, "spam words tend to be all too familiar". Laplace's rule does a similar job: by adding one to every count, it makes sure an unseen outcome keeps a small chance, and machine learning still uses it as Laplace smoothing.

Larger assumptions are built in as well. A network for pictures is built to look for the same patterns in every part of the frame, so a cat is a cat wherever it sits. Such built-in assumptions are called a model's inductive bias. With plenty of data they matter less and less; with little, they decide much of the answer. So AI researchers still argue over how much human knowledge to build in.

3 more questions from this passage

Extraordinary claims

How much evidence should it take to change your mind? Bayes' rule says it depends on where you started. In 2011 the psychologist Daryl Bem published experiments that seemed to show people sensing the future. Eric-Jan Wagenmakers and three colleagues replied with a worked example. Given everything else known, a sensible prior for such a power sits very close to zero; for illustration they took one chance in a hundred billion billion.

Now suppose a flawless experiment produced results 19 times more likely if the power were real than if it were not. That is strong evidence, and the rule multiplies the prior by about 19. The result is 19 chances in a hundred billion billion, still tiny. Laplace had said it first: the stranger a claim, the stronger the evidence it needs. Carl Sagan later made it a slogan: extraordinary claims require extraordinary evidence.

2 more questions from this passage

28 questions came out of this reading. Answer them out loud on your phone, and EdenMind schedules each one for the day you’re about to forget it.

Add to my practice