· turning point
Backpropagation
Rumelhart, Hinton and Williams show that multi-layer networks can learn internal representations by propagating errors backwards; Perceptrons is answered.
what had to happen · 7 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 5
- 1949Cells that fire together wire together
- 1958The perceptron learns
- 1960ADALINE and the least-mean-squares ruledirect
- 1969Perceptronsdirect
- 1970Reverse-mode automatic differentiationdirect
W1 · The first winter · 1
The paper is three pages long and appeared in Nature on 9 October 1986. Its claim is in the title: a network with hidden layers can learn representations, and the way to train it is to compute the error at the output and pass it backwards, layer by layer, using the chain rule, so that every weight in the network receives its share of the blame. The gradient that results is the direction to move, and moving a little at a time is gradient descent. David Rumelhart had worked it out in 1982; Geoffrey Hinton and Ronald Williams helped him make it convincing.
The mathematics was not new. Linnainmaa had the algorithm in 1970 and Werbos had applied it to networks in 1974. What the 1986 paper did was demonstrate, on problems people cared about, that the hidden units learned something interpretable, that the network solved XOR and family-tree relationships and the encoder problems Minsky and Papert had used as evidence of hopelessness. It arrived in the two-volume Parallel Distributed Processing books the same year, and a generation of researchers learned it from there.
Everything in the deep-learning era is trained this way. Convolutional networks, LSTMs, transformers, diffusion models: the architectures differ, and the loss functions differ, but the weights are always set by backpropagation and some variant of gradient descent. Rumelhart died in 2011. Hinton's 2024 Nobel citation begins with this paper.
what it led to · 96 events downstream, through 2026
Built on it directly:
- 1987NETtalk learns to read aloudII
- 1987The first NIPS conferenceII
- 1989ALVINN drives a van with a neural networkW2
- 1989The universal approximation theoremW2
- 1989LeNet reads handwritten postcodesW2
- 1990Finding structure in timeW2
- 1991The vanishing gradient problemW2
- 1992TD-Gammon reaches world-class backgammonW2
- 2003A neural probabilistic language modelIII
- 2006Deep belief networks and the word 'deep'III
- 2014AdamIV
- 2019The Turing Award goes to deep learningV
- 2024The Nobel Prizes go to neural networksVI
And, through them, by era:
III · Statistics and data · 7
IV · Deep learning · 17
- 2012Google Brain's network discovers cats
- 2012Dropout
- 2012AlexNet wins ImageNet
- 2013Word2vec
- 2013Deep Q-networks play Atari
- 2014Google buys DeepMind
- 2014Generative adversarial networks
- 2014Attention
- 2014Sequence to sequence learning
- 2015Batch normalisation
- 2015TensorFlow is open-sourced
- 2015Residual networks
- 2015OpenAI is founded
- 2016AlphaGo beats Lee Sedol
- 2016Google reveals the TPU
- 2016WaveNet
- 2016Google Translate goes neural
V · Transformers · 17
- 2017Attention is all you need
- 2017Deep reinforcement learning from human preferences
- 2017AlphaGo Zero learns from nothing
- 2018GPT: generative pre-training
- 2018BERT
- 2018AlphaFold enters the protein-folding contest
- 2019GPT-2 and the model too dangerous to release
- 2019The bitter lesson
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Learning to summarise from human feedback
- 2020An image is worth 16×16 words
- 2020AlphaFold 2 solves protein structure prediction
- 2021CLIP and DALL·E
- 2021On the dangers of stochastic parrots
- 2021Anthropic is founded
- 2021GitHub Copilot writes code
VI · Everyone · 28
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra