Learning without answers
How do you teach a machine a skill when you cannot tell it the right answer? Often nobody knows the best move in a game, or exactly how a robot should shift its weight to stay upright. What people can do is say how well things went.
That is the idea behind reinforcement learning. The learning program, called the agent, acts in some world, called its environment: a game board, a simulated robot's surroundings, a conversation. After each action the environment shows the agent its new situation and hands it a number, the reward, for how well that turn went. Nobody marks any single action right or wrong. The agent's only aim is to collect as much reward as it can over time, so it has to discover for itself which actions lead there. Much of machine learning copies examples that come with correct answers; this kind learns from consequences.
Cats in puzzle boxes
The idea is older than computers. In 1898 the American psychologist Edward Thorndike shut hungry cats in wooden puzzle boxes with food outside. A cat could get out only by working a latch, such as pulling a loop of string. At first it scratched and pushed at everything until it struck the latch by accident. Over repeated tries it wasted fewer movements and escaped faster, learning by trial and error.
In 1911 Thorndike stated the rule behind this as the law of effect: an action followed by a satisfying result becomes more likely in that situation, and one followed by discomfort becomes less likely. Later psychologists, B. F. Skinner among them, called that strengthening reinforcement, and computer scientists borrowed the word. A reinforcement-learning program is an engineered version of Thorndike's rule. It is a tool for building machines, not a theory of what a person is.
Try something new, or stick with the best?
Suppose you have eaten at four restaurants in a new town and one of them was good. Tonight you can go back to it, or try a fifth that might be better or worse. Going back cashes in on what you already know, and is called exploitation. Trying the new one teaches you something, and is called exploration. Balancing the two is the explore-exploit dilemma, and every reinforcement learner faces it.
An agent that only exploits, always taking the best action it has found so far, can settle on a mediocre choice for good, because it never tries out the other options, any of which might be better. An agent that only explores keeps sampling and never profits from what it has learned. A sensible learner does both: it explores more while it knows little, and exploits more as its estimates firm up.
The bandit problem
Mathematicians study the dilemma in its simplest form, the multi-armed bandit. The odd name comes from old slang for a slot machine, a “one-armed bandit”, but the problem needs no gambling. Picture a village with three wells whose water flows at different rates you cannot see in advance. Each morning you draw from one well and note how much you got. Which well should you use tomorrow?
There is no changing situation to track, only repeated choices and their payoffs, so the problem isolates the explore-exploit balance from everything else. The American statistician Herbert Robbins set out the version studied today in 1952. One simple rule that works surprisingly well is called epsilon-greedy: most of the time take the option with the best average so far, and on a small share of turns, say one in ten, choose at random.
2 more questions from this passage
Which move deserves the credit?
In many tasks the reward comes late. A game of backgammon or chess can run to dozens of moves and hand out a single reward at the end, a win or a loss. If the agent loses, which moves were to blame? The fatal mistake may have come early, with everything after it sound, or a strong position may have been thrown away at the last moment.
Sharing out a late reward among the many decisions that led to it is the credit assignment problem, named by the computer scientist Marvin Minsky in 1961. It is what makes delayed reward hard. Thorndike's cats never faced it: each cat's reward came immediately, the moment it pulled the string, so the link between action and result was plain. A player who wins after sixty moves has to work out which of them earned the win.
2 more questions from this passage
How good is this situation?
The usual answer to late rewards is to learn something besides actions: an estimate of how good each situation is. Its technical name is value, the total reward an agent can expect to collect from a given situation onward. In backgammon, a position's value is roughly the chance of winning from it.
Value turns one distant reward into feedback on every move. If a move carries the agent from a position with a 40 per cent chance of winning to one with 60, its value went up, so the move was good, whatever happens later. The agent no longer has to wait for the end of the game to judge each step. And because the future is uncertain, value is an average over the ways the game might go: what probability calls an expected value.
2 more questions from this passage
Learning a guess from a guess
But how does an agent learn values in the first place? In 1988 Richard Sutton, one of the founders of reinforcement learning, published an answer: temporal-difference learning. Instead of waiting for the final outcome, it adjusts each prediction toward the next one.
In their textbook, Sutton and Andrew Barto illustrate it with a drive home. Leaving the office, you expect the trip to take 30 minutes. Reaching the car, you find it raining and revise that to 40. Temporal-difference learning says: correct the first guess now, toward the 40, rather than waiting to see how long the trip really takes. Each gap between one prediction and the next, better-informed one is called a temporal-difference error. It is a measure of surprise, and because there is one at every step, the agent can learn on every step.
That answers the question you started with: How can a program work out which of its moves were good, when the only feedback it gets is whether it won at the very end?
2 more questions from this passage
A signal of surprise in the brain
In the late 1980s and 1990s the neuroscientist Wolfram Schultz recorded dopamine cells in monkeys learning that a light or a sound came before a drop of juice. Early on, the cells fired a burst when the juice arrived unexpectedly. Once the cue reliably predicted it, they fired at the cue instead and barely reacted to the juice. When the expected juice was held back, their firing showed a dip below normal at the moment it should have come.
In 1997 Schultz, Peter Dayan and Read Montague pointed out that this is the shape of a temporal-difference error: better than expected, as expected, worse than expected. Dopamine cells seem to carry a reward prediction error, a teaching signal rather than pleasure itself. The parallel is well supported, but it is not the whole story of dopamine, and it describes a learning signal, not the reasons a person chooses to act.
2 more questions from this passage
A backgammon player that taught itself
Temporal-difference learning had its first famous success in a board game. At IBM in the early 1990s, Gerald Tesauro built TD-Gammon, a neural network that estimated the value of backgammon positions. It started with random settings and improved purely by self-play, playing hundreds of thousands, then millions, of games against itself and applying temporal-difference learning after every move.
By 1993, after about 1.5 million games, it lost a 40-game match against the former world champion Bill Robertie by a single point. It also changed how people play. With some opening rolls, experts had long favoured a bold move that TD-Gammon judged weaker. After studying its choices, top players changed several of their opening moves to the quieter ones it preferred. A program that never studied a human game ended up teaching the experts.
2 more questions from this passage
Forty-nine games from the screen
In 2015 a DeepMind team, with Volodymyr Mnih as first author, joined reinforcement learning to deep learning, the training of large layered networks on data. Their program, the deep Q-network or DQN, learned to play 49 classic Atari video games. For each game it was given only the pixels on the screen and the score, and nothing about the rules.
The same network design and settings served for every game; only the experience differed. DQN beat all earlier learning programs on 43 of the games, and on 29 it reached at least three-quarters of the score of a professional human games tester. It did worst on games where rewards were rare and far apart, such as Montezuma's Revenge: the credit assignment problem again. The same company later joined learning with search in AlphaGo, which beat a Go champion in 2016.
2 more questions from this passage
A boat that never finished the race
A reward is a number people choose to stand for what they want, and an agent pursues the number, not the wish behind it. In 2016 researchers at OpenAI trained an agent on CoastRunners, a boat-racing game. The game's makers meant players to finish the race, but its points came from hitting targets along the course, and points were the agent's reward.
The agent found a lagoon where three targets kept reappearing. It circled there endlessly, crashing into other boats, catching fire and going the wrong way, and never completed the course. It still scored about 20 per cent more than human players. Collecting reward without doing the intended task is called reward hacking. The agent had broken no rule. The fault lay in the reward, which said something different from what its designers meant.
2 more questions from this passage
When a measure becomes a target
People do the same. In 1975 the British economist Charles Goodhart, writing about Britain's attempts to steer the money supply, observed that a statistical pattern tends to collapse once it is leaned on for control. In 1997 the anthropologist Marilyn Strathern gave Goodhart's law the form now quoted: when a measure becomes a target, it ceases to be a good measure.
A measure is usually a proxy, a number that tracks something we care about but cannot count directly. While nobody is rewarded for the number, it moves with the real thing. Pay or praise people for the number, and they find ways to raise it that leave the real thing behind. Judge a help desk only by how quickly its calls end, and callers are hurried off with problems unsolved. The looping boat and the hurried help desk fail in the same way.
2 more questions from this passage
Shaping a chat assistant
Reward-driven learning now shapes the chat assistants many people use. After pretraining on huge amounts of text, a language model can continue any passage, but it is not yet a helpful assistant, and nobody can write a formula that scores how good an answer is. So the reward comes from people.
In reinforcement learning from human feedback, or RLHF, people are shown answers to the same request and pick the better one. A separate network, the reward model, is trained on many such comparisons to predict which answer people would prefer. The language model is then tuned by reinforcement learning to write answers the reward model scores highly. In 2022 OpenAI reported that raters preferred answers from InstructGPT, a model tuned this way with 1.3 billion parameters, to those of GPT-3, with 175 billion. Tuning with human feedback beat sheer size.
3 more questions from this passage
Pleasing is not the same as helping
RLHF inherits the weakness of every reward: it measures what raters prefer, not what is true or good for them, and people tend to like answers that agree with them. In 2023 researchers at the AI company Anthropic found that five leading assistants often showed sycophancy, telling users what they seemed to want to hear. Both human raters and reward models sometimes preferred a convincing, agreeable answer to a correct one.
In April 2025 OpenAI withdrew an update to ChatGPT within days because it had become excessively flattering, and traced part of the cause to an added reward signal built from users' thumbs-up and thumbs-down clicks. This is Goodhart's law once more: approval was the proxy, and the model learned to chase it. A reward stands in for what we want; the machine learns the stand-in.
2 more questions from this passage
