Rectified linear units
Nair and Hinton replace the sigmoid with max(0, x); the gradient no longer vanishes through active units and deep networks train several times faster.
what had to happen · 13 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 5
W1 · The first winter · 1
II · Connection · 3
W2 · The second winter · 2
- 1990Finding structure in time
- 1991The vanishing gradient problemdirect
III · Statistics and data · 1
- 2006Deep belief networks and the word 'deep'direct
For two decades the standard artificial neuron squashed its input through a sigmoid, a smooth S-curve that saturates at both ends. The saturation is where the gradient dies: a unit that is strongly on or strongly off passes almost no error signal back. Vinod Nair and Geoffrey Hinton's ICML paper in June 2010 tried the simplest possible alternative, output the input if it is positive and zero otherwise, and found that it worked better in restricted Boltzmann machines. Xavier Glorot, Antoine Bordes and Yoshua Bengio showed the next year that it worked better in deep supervised networks too, with no pretraining needed.
The rectifier's derivative is one for any active unit, so the gradient passes through undiminished however deep the network, and half the units are exactly zero at any time, which is cheap and sparse. It is a McCulloch–Pitts threshold with a linear ramp above it, and the field had walked past it for sixty years because it was not differentiable at zero and did not look like a neuron.
AlexNet used it in 2012 and reported training six times faster than with sigmoids. Almost every network since has used it or a close relative. Along with GPUs, data and dropout, it is one of the four ingredients that made 2012 possible.
what it led to · 70 events downstream, through 2026
Built on it directly:
- 2012AlexNet wins ImageNetIV
And, through them, by era:
IV · Deep learning · 9
V · Transformers · 17
- 2017Attention is all you need
- 2017Deep reinforcement learning from human preferences
- 2017AlphaGo Zero learns from nothing
- 2018GPT: generative pre-training
- 2018BERT
- 2018AlphaFold enters the protein-folding contest
- 2019GPT-2 and the model too dangerous to release
- 2019The bitter lesson
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Learning to summarise from human feedback
- 2020An image is worth 16×16 words
- 2020AlphaFold 2 solves protein structure prediction
- 2021CLIP and DALL·E
- 2021On the dangers of stochastic parrots
- 2021Anthropic is founded
- 2021GitHub Copilot writes code
VI · Everyone · 29
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra