Deep learning moves to GPUs
Raina, Madhavan and Ng train deep belief networks on graphics cards seventy times faster than on CPUs; the hardware and the method find each other.
what had to happen · 14 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 1
II · Connection · 3
III · Statistics and data · 3
- 1999The first GPU
- 2006Deep belief networks and the word 'deep'direct
- 2007CUDAdirect
Andrew Ng's group at Stanford wanted to train unsupervised models with a hundred million parameters, and on the CPUs of 2008 that took weeks. Their paper at ICML in June 2009 reported what happened when they wrote the training code in CUDA for an NVIDIA GTX 280. Deep belief networks trained up to seventy times faster; sparse coding, fifteen times. Models that had been out of reach were trained in a day.
The arithmetic was not surprising to anyone who had looked. A graphics card of the time had 240 cores and a memory bandwidth ten times that of a CPU, and neural-network training is almost entirely dense linear algebra. What the paper did was demonstrate it on the models people were actually trying to build, and publish the speed-ups in a venue the field read.
Within three years the practice was universal. Dan Cireşan's group in Switzerland set records on MNIST and traffic signs with GPU-trained networks in 2010 and 2011; Alex Krizhevsky trained AlexNet on two consumer cards in 2012. The lesson Ng drew, that scale in compute and data mattered more than cleverness in models, took him to Google Brain in 2011 and its billion-parameter experiments.
what it led to · 49 events downstream, through 2026
Built on it directly:
And, through them, by era:
IV · Deep learning · 1
V · Transformers · 6
VI · Everyone · 27
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra