Residual networks
He, Zhang, Ren and Sun add skip connections so each layer learns a correction to its input; 152-layer networks train easily and beat humans on ImageNet.
what had to happen · 24 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 3
III · Statistics and data · 6
IV · Deep learning · 3
- 2012Dropout
- 2012AlexNet wins ImageNetdirect
- 2015Batch normalisationdirect
Deeper networks should be at least as good as shallower ones, since the extra layers could simply pass their input through. In practice, by 2015, networks past about twenty layers got worse, not from overfitting but because the optimiser could not find the pass-through solution. Kaiming He's group at Microsoft Research Asia made the pass-through the default. Each block of layers computes a change to its input and adds it, so a block that learns nothing does nothing, and the gradient has a clear path back through the identity connections however deep the stack.
The paper, posted on 10 December 2015, trained a 152-layer network on ImageNet and won the 2015 challenge with 3.57 percent top-five error, below the estimated human rate, along with the detection and segmentation competitions. A 1,000-layer version trained without difficulty.
Residual connections are now in everything. The transformer of 2017 is a stack of residual blocks, and the reason language models can be a hundred layers deep is the identity path He introduced for pictures. The paper is among the most cited in computer science, and its central idea, that a layer should learn a small correction rather than a whole new representation, is the last piece of the deep-learning architecture that was needed.
what it led to · 57 events downstream, through 2026
Built on it directly:
- 2017Attention is all you needV
- 2018AlphaFold enters the protein-folding contestV
- 2020An image is worth 16×16 wordsV
And, through them, by era:
V · Transformers · 11
- 2018GPT: generative pre-training
- 2018BERT
- 2019GPT-2 and the model too dangerous to release
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Learning to summarise from human feedback
- 2020AlphaFold 2 solves protein structure prediction
- 2021CLIP and DALL·E
- 2021On the dangers of stochastic parrots
- 2021Anthropic is founded
- 2021GitHub Copilot writes code
VI · Everyone · 29
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra