Batch normalisation
Ioffe and Szegedy normalise the activations inside a network during training; deep networks train in a fraction of the steps and much deeper stacks become practical.
what had to happen · 23 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 3
III · Statistics and data · 6
IV · Deep learning · 2
- 2012Dropout
- 2012AlexNet wins ImageNetdirect
As a deep network trains, each layer's inputs shift because the layers below it are changing, and every layer is chasing a moving target. Sergey Ioffe and Christian Szegedy's fix, posted on 11 February 2015, was to standardise each layer's inputs over the current mini-batch, subtracting the mean and dividing by the standard deviation, and then to let the network learn a scale and shift of its own. The paper's explanation, reducing internal covariate shift, was later disputed; the effect was not. Networks trained with batch normalisation reached the same accuracy in fourteen times fewer steps, tolerated much higher learning rates, and needed less dropout.
The immediate result was that Google's Inception network beat human-level accuracy on ImageNet, at 4.8 percent top-five error, in the same paper. The longer result was depth. ResNet later that year stacked 152 layers, and could not have been trained without normalisation inside it.
The technique's dependence on the batch made it awkward for sequences, and the transformer of 2017 used layer normalisation, which standardises across a layer's units for each example instead. But the principle, keep the activations well scaled and the gradients will behave, is one of the reasons networks of the 2020s can be a thousand layers deep.
what it led to · 58 events downstream, through 2026
Built on it directly:
- 2015Residual networksIV
And, through them, by era:
V · Transformers · 14
- 2017Attention is all you need
- 2018GPT: generative pre-training
- 2018BERT
- 2018AlphaFold enters the protein-folding contest
- 2019GPT-2 and the model too dangerous to release
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Learning to summarise from human feedback
- 2020An image is worth 16×16 words
- 2020AlphaFold 2 solves protein structure prediction
- 2021CLIP and DALL·E
- 2021On the dangers of stochastic parrots
- 2021Anthropic is founded
- 2021GitHub Copilot writes code
VI · Everyone · 29
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra