Skip to content
The Shape of Intelligence

TD-Gammon reaches world-class backgammon

Gerald Tesauro's network learns backgammon by playing itself with temporal-difference learning and reaches the level of the best humans, changing how they play.

category
model
significance
4 of 5
people
Gerald Tesauro
organisations
IBM

what had to happen · 12 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

TD-Gammon was a network with one hidden layer of forty units, later eighty, that took a description of a backgammon position and predicted who would win. Gerald Tesauro at IBM trained it by having it play itself, hundreds of thousands of games and eventually over a million, adjusting its predictions after every move with Sutton's temporal-difference rule. It was given no strategy and no expert games. By 1992 it played at a strong level; by 1995 it was as good as the best humans in the world, and a few of its opening moves, which contradicted a century of received wisdom, were adopted by them.

It was the first program to reach the top of a serious game by learning rather than by search, and it did so in the middle of the second winter, on hardware that would fit in a wristwatch today. Tesauro's papers were read carefully by a small community and ignored by the mainstream, which regarded backgammon's dice as a special case.

They were the special case that generalised. AlphaGo Zero in 2017 is TD-Gammon's recipe with a deep network and tree search: self-play, a value function learned from outcomes, no human data. Tesauro's forty hidden units turned out to have been the right idea waiting for thirty years of Moore's law.

what it led to · 45 events downstream, through 2026

Built on it directly:

  1. 2010DeepMind is foundedIII
  2. 2013Deep Q-networks play AtariIV
  3. 2016AlphaGo beats Lee SedolIV
  4. 2017AlphaGo Zero learns from nothingV

And, through them, by era:

sources · 2

See this era in the exhibition →Back to the timeline