Skip to content
The Shape of Intelligence

AlphaGo Zero learns from nothing

A new version starts from random play with no human games at all, and after three days beats the AlphaGo that beat Lee Sedol 100 games to 0.

category
model
significance
4 of 5
people
David Silver, Julian Schrittwieser, Demis Hassabis
organisations
DeepMind

what had to happen · 31 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

The original AlphaGo had started from thirty million moves by human experts. The version published in Nature on 18 October 2017 started from nothing. A single network, given only the rules, played itself, and at every move used tree search to improve on its own raw guess; the improved move became the training target, and the game's outcome trained the value. After three days and 4.9 million games it beat the Lee Sedol version 100 games to nil, and after forty days it beat the version that had defeated the world number one, Ke Jie, earlier that year.

It also rediscovered Go. Its self-play games passed through the standard human openings and beyond them, and some joseki that players had used for centuries it abandoned as inferior. Its successor AlphaZero, two months later, learned chess and shogi the same way in hours and beat the strongest engines.

AlphaGo Zero is the purest demonstration on this timeline of Sutton's bitter lesson before he wrote it: remove the human knowledge, add compute and self-play, and the result is stronger. It is also TD-Gammon's method vindicated twenty-five years on, and the origin of the search-plus-learning loop that reasoning models revived for language in 2024.

what it led to · 9 events downstream, through 2026

Built on it directly:

  1. 2024o1 and reasoning modelsVI

And, through them, by era:

sources · 1

See this era in the exhibition →Back to the timeline