The Boltzmann machine
Ackley, Hinton and Sejnowski add noise and hidden units to the Hopfield network and derive a learning rule, the first for a network with hidden layers.
what had to happen · 3 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
A Hopfield network could store patterns but could not learn new features, because every neuron was either an input or an output. Geoffrey Hinton and Terrence Sejnowski, with David Ackley, added hidden units, neurons connected only to other neurons, and made every unit stochastic, flipping on and off with a probability set by its energy and a temperature, as in the statistical mechanics of Ludwig Boltzmann. The network then had a probability distribution over states, and learning meant changing the weights so that the distribution over the visible units matched the data.
The learning rule they derived in 1985 is elegant and slow. Run the machine clamped to the data and measure how often pairs of units are on together; run it free and measure again; move each weight in proportion to the difference. It was the first working algorithm for training hidden units, a year before backpropagation, and the first generative model in the modern sense: a network whose job was to reproduce the statistics of its inputs.
Its restricted form, with connections only between visible and hidden layers, was what Hinton used in 2006 to train deep networks one layer at a time and end the second winter. The Nobel citation of 2024 names it alongside Hopfield's memory.
what it led to · 79 events downstream, through 2026
Built on it directly:
- 2006Deep belief networks and the word 'deep'III
- 2014Generative adversarial networksIV
- 2015Diffusion modelsIV
And, through them, by era:
III · Statistics and data · 2
IV · Deep learning · 12
V · Transformers · 19
- 2017Attention is all you need
- 2017Deep reinforcement learning from human preferences
- 2017AlphaGo Zero learns from nothing
- 2018GPT: generative pre-training
- 2018BERT
- 2018AlphaFold enters the protein-folding contest
- 2019GPT-2 and the model too dangerous to release
- 2019The bitter lesson
- 2019The Turing Award goes to deep learning
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Denoising diffusion probabilistic models
- 2020Learning to summarise from human feedback
- 2020An image is worth 16×16 words
- 2020AlphaFold 2 solves protein structure prediction
- 2021CLIP and DALL·E
- 2021On the dangers of stochastic parrots
- 2021Anthropic is founded
- 2021GitHub Copilot writes code
VI · Everyone · 29
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra