Psychology

When famous findings fall apart

Some of the best-known stories about the mind met the simplest test in science: doing the study again.

  • 9min read
  • 11min listen
  • 30questions
A paper stone arch, most stones set firm, with a few pink stones near its top slipping loose and one falling.

When famous findings fall apart

0:00 / 11:06

A question to hold while you read

Why do the effects in famous psychology studies so often shrink when someone runs the same study again?

Doing it again

How do scientists know a finding is real and not a fluke? One study, however careful, can mislead. Its participants might have been unusual, a measurement might have slipped, or chance alone might have produced a striking pattern. So science leans on replication: an independent team runs the same study again, following the same method with new people, and sees whether the result comes back.

A result that returns time after time, in different labs and countries, earns trust. A result that appears once and never again was probably luck or error. That is why replication is called the backbone of science. For decades, though, psychologists rarely did it. Journals wanted new discoveries, and repeating someone else's experiment brought little credit. Around 2011 a run of shocks made the field look back at its most famous findings and ask how many would survive being run again.

An experiment that saw the future

In 2011 the social psychologist Daryl Bem published nine experiments in one of the field's most respected journals. They claimed evidence for precognition: sensing events before they happen. In one, students saw a list of words, tried to recall them, and only afterwards practised some of the words. They had recalled the practised words better, as if the later practice had reached back in time.

More than a thousand people took part, and eight of the nine experiments passed the standard statistical test. That was the alarm. Bem had used the ordinary methods of the field, and almost nobody believed his conclusion. If ordinary methods could prove the impossible, what else might they have proved? When three British researchers repeated the recall experiment three times and found nothing, the journal that had published Bem turned their paper away. It did not publish repeats of earlier studies.

What "significant" means

What was the standard test Bem passed? Most psychology studies compare two groups and ask whether a difference between them is real or just the luck of who happened to land in each group. The answer comes as a p-value: the chance of seeing a difference at least this large if the thing being tested had no effect at all.

By long convention, a p-value below 0.05, one chance in twenty, counts as statistically significant and fit to publish. The catch sits in that one in twenty. Test something with no effect twenty times and, on average, one of those tests will still come out significant. That lucky hit is a false positive: evidence for something that is not there. One in twenty sounds safe, but only if each study makes a single test that was planned in advance. Many studies did not.

Many roads to "significant"

In 2011 the psychologists Joseph Simmons, Leif Nelson and Uri Simonsohn showed how easily honest researchers could fool themselves. Every study involves small choices: which of two measures to report, whether to test ten more people per group, whether to adjust for sex, whether to drop one of three groups. They called this latitude researcher degrees of freedom, and trying choice after choice until one works is now known as p-hacking.

Each choice looks harmless. But trying several and reporting the one that works gives chance many more tickets to win. In computer simulations, combining just four such choices raised the false-positive rate from 5 per cent to about 61 per cent. To make the point, they ran a real experiment and "showed" that listening to the Beatles song When I'm Sixty-Four left people nearly a year and a half younger.

5 per centone planned test61 per centfour choices combined
How often a study of a non-existent effect comes out significant, in the 2011 simulations

The file drawer

Imagine twenty labs test the same idea, and the idea is wrong. Nineteen find nothing. One finds a significant result by chance. Journals, keen on clear and surprising discoveries, publish that one. The nineteen null results stay in a file drawer, unpublished, so readers see only the success. The psychologist Robert Rosenthal named this the file drawer problem in 1979.

The resulting skew in what reaches print is called publication bias, and incentives drive it. Careers are built on publications, and a surprising positive result is far easier to publish than a failure or a repeat of someone else's work. So the published record can fill up with findings that each look solid, while the failures that would put them in context are nowhere to be seen.

2 more questions from this passage

Small studies, big claims

Many classic studies tested only a few dozen people. That matters because of statistical power: the chance that a study will detect an effect that really exists. Small samples are noisy, so a small study of a modest real effect will usually miss it. With twenty people in each group, a study of a modest effect finds it less than a quarter of the time.

Small samples also distort the hits. When a small study does reach significance, it is often because chance pushed its result unusually high that time. So the effects that get published from small studies tend to come out larger than the true effect. Run again with a big sample, the same effect shrinks, or vanishes if it was never there. The noise that made the first result exciting is exactly what the second study averages away.

That answers the question you started with: Why do the effects in famous psychology studies so often shrink when someone runs the same study again?

2 more questions from this passage

One hundred studies, run again

How widespread was the problem? In 2015 the Open Science Collaboration, 270 researchers coordinated by Brian Nosek, published an answer in the journal Science. In the Reproducibility Project, they had repeated 100 studies published in 2008 in three leading psychology journals, using larger samples and, where possible, the original materials.

Of the originals, 97 had reported significant results. Of the repeats, only 36 per cent did. On average, the effects the repeats found were about half the size of the originals.

That did not mean the rest of psychology was false. A repeat can itself fail by chance, or copy the original imperfectly, and some effects were real but smaller than first claimed. What the project did show was that one significant finding in a good journal was far weaker evidence than most readers had assumed.

97 per centoriginals36 per centrepeats
Studies with significant results in the 2015 Reproducibility Project

3 more questions from this passage

The slow walk down the corridor

One of the most famous findings to be retested was about priming: the idea that a word or picture can nudge later behaviour without our noticing. In a 1996 study by John Bargh and his colleagues, students unscrambled sentences sprinkled with words linked to old age, such as Florida and wrinkle. As they left, they walked more slowly down the corridor than students who had seen neutral words.

In 2012 Stéphane Doyen and colleagues in Belgium repeated the study with more people, timing the walk with infrared sensors instead of a hand-held stopwatch. The slow walking disappeared. Then they led some of the experimenters to expect slow walking. Only then did it come back, and only with the experimenters who expected it. What the experimenters expected, not the words, may have produced the original effect.

3 more questions from this passage

A train wreck looming

The slow-walking study had a famous admirer. In his 2011 book Thinking, Fast and Slow, Daniel Kahneman presented priming research as a showcase of System 1, the fast, automatic mind that works beneath deliberate thought. In September 2012, as failed repeats piled up, he wrote an open letter to priming researchers: "I see a train wreck looming." He urged them to form a chain of labs, each rerunning a neighbour's study.

In 2017 he went further. Answering a blog analysis of his book's priming chapter, he wrote that he had "placed too much faith in underpowered studies". The irony was that his first paper with Amos Tversky, in 1971, had warned that researchers trust small samples too much. Surprising, elegant findings make good stories, and System 1 believes good stories easily. That pull works on scientists and readers alike.

19952000200520102015slow-walking studyKahneman's bookfailed repeat, open letterKahneman's admission
How behavioural priming went from showcase to warning

3 more questions from this passage

Willpower that runs out

Another celebrated idea was ego depletion: that self-control draws on a limited store, so using it up on one task leaves less for the next. In a 1998 study led by Roy Baumeister, hungry students who had to eat radishes while resisting fresh chocolate cookies later gave up sooner on an unsolvable puzzle than students who had been allowed the sweets. Hundreds of studies followed, and a 2010 review of 83 of them found a medium-sized effect.

Yet that review could count only published studies. In 2016, 23 labs ran one agreed procedure with 2,141 participants, having registered their plan and committed to report whatever they found. The effect came out close to zero. Defenders objected to the task used, so in 2021 a project of 36 labs, led by researchers who had long studied the effect, tried again. It also found almost nothing.

0.622010 review, 83 studies0.042016 test, 23 labs0.062021 test, 36 labs
Size of the ego-depletion effect (a standard measure, Cohen's d)

2 more questions from this passage

The prison study

The Stanford Prison Experiment of 1971 is among the most famous studies in psychology. Philip Zimbardo split student volunteers into guards and prisoners in a mock jail, and the study was stopped after six days as the guards grew abusive. For decades it was told as proof that ordinary people turn cruel when handed a role.

It was never repeated in the same form. When the psychologists Alex Haslam and Stephen Reicher ran a partly similar prison study for BBC television, filmed in late 2001, the guards never settled into their role, and the prisoners eventually overpowered them. In 2018 the French researcher Thibault Le Texier, working from the Stanford archives, showed that the guards were briefed on how to behave. That points to demand characteristics: cues that tell participants what the experimenter wants to see.

1 more question from this passage

Deciding before the data

The most direct fix is to decide before looking. In preregistration, researchers post a time-stamped plan online before collecting any data: the hypothesis, the sample size and the exact analysis. Choices made after seeing the results then stand out as departures from the plan.

Registered Reports, launched at the journal Cortex in 2013 by the neuroscientist Chris Chambers, go a step further. Reviewers judge the question and the method before the study is run, and a sound plan is accepted before the results are known. A null result is then as publishable as a striking one. The difference shows. A 2021 comparison found that the first hypothesis was supported in 96 per cent of standard psychology papers, but in only 44 per cent of Registered Reports.

3 more questions from this passage

Many labs at once

Another fix is scale. In the Many Labs project, published in 2014, labs around the world ran the same 13 classic and newer effects in 36 samples totalling 6,344 people, then pooled the results. Ten of the 13 replicated, and one had only weak support. Two did not replicate, both of them priming effects: seeing a national flag did not make people more conservative, and images of money did not make them defend the social system.

One of the ten, anchoring, came back even stronger than first reported. That is the pull an arbitrary first number exerts on a later estimate, and Kahneman had helped establish it. Projects like this, with data and methods open for anyone to check, show that the crisis did not prove psychology empty. It was science correcting itself, keeping what survives and dropping what does not.

10replicated1weak support2did not replicate
What happened to 13 effects in the 2014 Many Labs project

2 more questions from this passage

An older habit of checking

Trusting a claim only when others confirm it is an old idea. Muslim scholars of hadith, the reports of the Prophet Muhammad's words and deeds, judged each report partly by its isnād, the chain of people who had passed it on. One test was a narrator's precision, called ḍabṭ. Critics compared what a narrator reported from a teacher with what that teacher's other students reported. A narrator whose peers usually matched him was trusted; one who often contradicted them was not.

Critics also searched for corroboration, support for a report through other chains, which could strengthen one whose own chain was weak. A report carried by so many independent chains that they could not have agreed on a falsehood, called mutawātir, was beyond doubt. The methods are not statistics, and the scholars were weighing people, not experiments. But the instinct is shared: trust grows with independent confirmation.

2 more questions from this passage

30 questions came out of this reading. Answer them out loud on your phone, and EdenMind schedules each one for the day you’re about to forget it.

Add to my practice