Temporal-difference learning
Richard Sutton formalises learning from the difference between successive predictions, the method inside Samuel's checkers player, TD-Gammon and AlphaGo.
what had to happen · 3 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
I · Foundations · 3
Suppose you are predicting how a game will end, and each move you make a new prediction. Ordinary supervised learning would wait for the result and then correct every prediction against it. Richard Sutton's 1988 paper argues that you should instead correct each prediction against the next one, moment by moment, before the outcome is known. The error signal is the temporal difference, and learning from it is cheaper, works online, and, he proved, converges to the right answer for problems with the Markov property.
Sutton had been developing the idea with Andrew Barto since the late 1970s, out of the psychology of animal learning and the trial-and-error machines of the 1950s. The paper made it a general method and connected it to dynamic programming; Chris Watkins's Q-learning the next year extended it from prediction to control, and their textbook of 1998 defined reinforcement learning as a field.
Almost every system on this timeline that learned by playing runs on TD. Tesauro's TD-Gammon in 1992, DeepMind's Atari player in 2013 and AlphaGo in 2016 all learn value functions by bootstrapping one prediction from the next. Dopamine neurons in the brain, it was found in the 1990s, appear to signal exactly Sutton's error.
what it led to · 47 events downstream, through 2026
Built on it directly:
And, through them, by era:
III · Statistics and data · 1
IV · Deep learning · 4
V · Transformers · 7
VI · Everyone · 19
- 2022InstructGPT
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra