Skip to content
The Shape of Intelligence

Temporal-difference learning

Richard Sutton formalises learning from the difference between successive predictions, the method inside Samuel's checkers player, TD-Gammon and AlphaGo.

category
theory
significance
4 of 5
people
Richard Sutton
organisations
GTE Laboratories, University of Massachusetts Amherst

what had to happen · 3 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

Suppose you are predicting how a game will end, and each move you make a new prediction. Ordinary supervised learning would wait for the result and then correct every prediction against it. Richard Sutton's 1988 paper argues that you should instead correct each prediction against the next one, moment by moment, before the outcome is known. The error signal is the temporal difference, and learning from it is cheaper, works online, and, he proved, converges to the right answer for problems with the Markov property.

Sutton had been developing the idea with Andrew Barto since the late 1970s, out of the psychology of animal learning and the trial-and-error machines of the 1950s. The paper made it a general method and connected it to dynamic programming; Chris Watkins's Q-learning the next year extended it from prediction to control, and their textbook of 1998 defined reinforcement learning as a field.

Almost every system on this timeline that learned by playing runs on TD. Tesauro's TD-Gammon in 1992, DeepMind's Atari player in 2013 and AlphaGo in 2016 all learn value functions by bootstrapping one prediction from the next. Dopamine neurons in the brain, it was found in the 1990s, appear to signal exactly Sutton's error.

what it led to · 47 events downstream, through 2026

Built on it directly:

  1. 1989Q-learningW2
  2. 1992TD-Gammon reaches world-class backgammonW2

And, through them, by era:

sources · 1

See this era in the exhibition →Back to the timeline