The vanishing gradient problem
Sepp Hochreiter's diploma thesis shows why deep and recurrent networks fail to learn: error signals shrink exponentially as they travel back through layers.
what had to happen · 9 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 5
W1 · The first winter · 1
II · Connection · 1
- 1986Backpropagationdirect
W2 · The second winter · 1
- 1990Finding structure in timedirect
Backpropagation worked for networks with one or two hidden layers and stalled for deeper ones, and for recurrent networks it could not learn dependencies more than a few steps apart. Nobody knew exactly why until Sepp Hochreiter's 1991 diploma thesis, written under Jürgen Schmidhuber in Munich. The gradient that reaches an early layer is a product of the derivatives of every layer after it, and with sigmoid units those derivatives are less than one. Multiplied together across many layers or time steps, the signal shrinks towards zero; with other weight settings it explodes. Either way the early layers receive nothing useful to learn from.
The thesis was in German and read by few, but the analysis explained a decade of failures and set the agenda for fixing them. Hochreiter and Schmidhuber's LSTM of 1997 built a memory cell whose gradient could pass through unchanged. Rectified linear units, careful initialisation, batch normalisation and residual connections, all of them the standard equipment of the deep-learning era, are answers to the problem stated here.
The thesis is also why "deep" was not a compliment for another fifteen years: until the fixes arrived, depth was the thing that stopped networks learning.
what it led to · 75 events downstream, through 2026
Built on it directly:
- 1997Long short-term memoryIII
- 2010Rectified linear unitsIII
- 2015Batch normalisationIV
And, through them, by era:
IV · Deep learning · 12
V · Transformers · 17
- 2017Attention is all you need
- 2017Deep reinforcement learning from human preferences
- 2017AlphaGo Zero learns from nothing
- 2018GPT: generative pre-training
- 2018BERT
- 2018AlphaFold enters the protein-folding contest
- 2019GPT-2 and the model too dangerous to release
- 2019The bitter lesson
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Learning to summarise from human feedback
- 2020An image is worth 16×16 words
- 2020AlphaFold 2 solves protein structure prediction
- 2021CLIP and DALL·E
- 2021On the dangers of stochastic parrots
- 2021Anthropic is founded
- 2021GitHub Copilot writes code
VI · Everyone · 29
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra