Deep belief networks and the word 'deep'
Hinton, Osindero and Teh train a deep network one layer at a time as stacked Boltzmann machines and then fine-tune it; deep learning gets its name and its first results.
what had to happen · 10 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 5
W1 · The first winter · 1
II · Connection · 3
- 1982The Hopfield network
- 1985The Boltzmann machinedirect
- 1986Backpropagationdirect
By 2006 Geoffrey Hinton had spent two decades on neural networks through two winters, funded latterly by a Canadian institute that had decided to back unfashionable ideas. That summer two papers came out of his Toronto group. The first, in Neural Computation, showed how to train a deep network by treating each pair of layers as a restricted Boltzmann machine, training them greedily from the bottom up, and then fine-tuning the whole stack with backpropagation. The second, in Science, used the method to build deep autoencoders that compressed images and documents far better than the standard techniques.
The trick got around the vanishing gradient by not relying on it: the layers were already sensible before backpropagation started. Deep networks, which had been unstable and unfashionable, became trainable, and the phrase "deep learning" was adopted to describe them, partly as a rebrand for a field that had learned to avoid the words neural network.
The layer-wise pretraining was abandoned within a few years, once rectified units, better initialisation and GPUs made it unnecessary. What survived was the confidence. The 2006 papers convinced a small group of people that depth was the direction, and that group produced the results of 2012.
what it led to · 76 events downstream, through 2026
Built on it directly:
- 2009Deep learning moves to GPUsIII
- 2010Rectified linear unitsIII
- 2012Google Brain's network discovers catsIV
- 2012DropoutIV
- 2019The Turing Award goes to deep learningV
And, through them, by era:
IV · Deep learning · 11
V · Transformers · 17
- 2017Attention is all you need
- 2017Deep reinforcement learning from human preferences
- 2017AlphaGo Zero learns from nothing
- 2018GPT: generative pre-training
- 2018BERT
- 2018AlphaFold enters the protein-folding contest
- 2019GPT-2 and the model too dangerous to release
- 2019The bitter lesson
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Learning to summarise from human feedback
- 2020An image is worth 16×16 words
- 2020AlphaFold 2 solves protein structure prediction
- 2021CLIP and DALL·E
- 2021On the dangers of stochastic parrots
- 2021Anthropic is founded
- 2021GitHub Copilot writes code
VI · Everyone · 29
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra