TD-Gammon reaches world-class backgammon
Gerald Tesauro's network learns backgammon by playing itself with temporal-difference learning and reaches the level of the best humans, changing how they play.
what had to happen · 12 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 8
- 1948A mathematical theory of communication
- 1949Cells that fire together wire together
- 1950Programming a computer for playing chess
- 1958The perceptron learns
- 1959Samuel's checkers program coins 'machine learning'direct
- 1960ADALINE and the least-mean-squares rule
- 1969Perceptrons
- 1970Reverse-mode automatic differentiation
W1 · The first winter · 1
II · Connection · 1
- 1986Backpropagationdirect
W2 · The second winter · 1
- 1988Temporal-difference learningdirect
TD-Gammon was a network with one hidden layer of forty units, later eighty, that took a description of a backgammon position and predicted who would win. Gerald Tesauro at IBM trained it by having it play itself, hundreds of thousands of games and eventually over a million, adjusting its predictions after every move with Sutton's temporal-difference rule. It was given no strategy and no expert games. By 1992 it played at a strong level; by 1995 it was as good as the best humans in the world, and a few of its opening moves, which contradicted a century of received wisdom, were adopted by them.
It was the first program to reach the top of a serious game by learning rather than by search, and it did so in the middle of the second winter, on hardware that would fit in a wristwatch today. Tesauro's papers were read carefully by a small community and ignored by the mainstream, which regarded backgammon's dice as a special case.
They were the special case that generalised. AlphaGo Zero in 2017 is TD-Gammon's recipe with a deep network and tree search: self-play, a value function learned from outcomes, no human data. Tesauro's forty hidden units turned out to have been the right idea waiting for thirty years of Moore's law.
what it led to · 45 events downstream, through 2026
Built on it directly:
- 2010DeepMind is foundedIII
- 2013Deep Q-networks play AtariIV
- 2016AlphaGo beats Lee SedolIV
- 2017AlphaGo Zero learns from nothingV
And, through them, by era:
IV · Deep learning · 2
V · Transformers · 6
VI · Everyone · 19
- 2022InstructGPT
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra