Skip to content
The Shape of Intelligence

The bitter lesson

Richard Sutton's short essay argues that seventy years of AI show one thing: methods that use more computation beat methods that use more human knowledge, every time.

category
culture
significance
3 of 5
people
Richard Sutton
organisations
University of Alberta, DeepMind

what had to happen · 31 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

Richard Sutton, who had formalised temporal-difference learning in 1988, published a thousand-word essay on his website on 13 March 2019 that became the most quoted document of the scaling era. Its argument is historical. In chess, the researchers who encoded human knowledge lost to Deep Blue's search. In Go, they lost to self-play. In speech and vision, hand-built features lost to learning from data. Each time, the field's instinct was to build in what it knew, and each time a method that instead used more computation, general search and general learning, won, because computation gets cheaper by Moore's law and human knowledge does not.

The lesson is bitter, he wrote, because it is a loss for the researchers' own contributions: "the only thing that matters in the long run is the leveraging of computation." What should be built in is not knowledge but the capacity to acquire it.

The essay was published a year before the scaling laws quantified it and three years before ChatGPT made it obvious. It is cited by those who build ever-larger models as a justification and by their critics as the confession of a field that has stopped thinking. Either way, the decade since has not produced a counterexample.

what it led to · 0 events downstream

A leaf, for now. Nothing in the archive has built on it yet.

sources · 1

See this era in the exhibition →Back to the timeline