Skip to content
The Shape of Intelligence

· turning point

AlexNet wins ImageNet

Krizhevsky, Sutskever and Hinton's convolutional network, trained on two gaming GPUs, cuts the ImageNet error rate from 26% to 15%; the deep-learning era begins.

category
model
significance
5 of 5
people
Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton
organisations
University of Toronto

what had to happen · 22 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

The results of the 2012 ImageNet challenge were released on 30 September. The second-placed entry, from the University of Tokyo, had a top-five error rate of 26.2 percent, a typical increment on the previous year. The winner, a network called SuperVision from the University of Toronto, had 15.3 percent. Nothing in the history of the benchmark, or of the field, had moved that far in one step, and the vision community, which had mostly regarded neural networks as a relic, changed its mind within a year.

The network was Alex Krizhevsky's, with Ilya Sutskever and their supervisor Geoffrey Hinton. It was LeNet's design, eight layers deep, with 60 million parameters, rectified units, dropout, and data augmentation, trained for six days on two NVIDIA GTX 580 gaming cards in Krizhevsky's bedroom using CUDA code he had written himself. Every ingredient existed already; the combination, at that scale, had never been tried.

Within two years every entry in the challenge was a deep convolutional network, error rates fell below human level by 2015, and Google, Facebook, Microsoft and Baidu had bought or built deep-learning groups. Hinton, Krizhevsky and Sutskever's small company was acquired by Google in 2013. If the deep-learning era has a birthday, it is this one.

what it led to · 69 events downstream, through 2026

Built on it directly:

  1. 2013Deep Q-networks play AtariIV
  2. 2014Generative adversarial networksIV
  3. 2015Batch normalisationIV
  4. 2015Residual networksIV
  5. 2016AlphaGo beats Lee SedolIV
  6. 2016Google reveals the TPUIV
  7. 2016WaveNetIV
  8. 2019The bitter lessonV
  9. 2025Nvidia is worth four trillion dollarsVII

And, through them, by era:

IV · Deep learning · 2
  1. 2014Google buys DeepMind
  2. 2015OpenAI is founded
V · Transformers · 16
  1. 2017Attention is all you need
  2. 2017Deep reinforcement learning from human preferences
  3. 2017AlphaGo Zero learns from nothing
  4. 2018GPT: generative pre-training
  5. 2018BERT
  6. 2018AlphaFold enters the protein-folding contest
  7. 2019GPT-2 and the model too dangerous to release
  8. 2020Scaling laws for neural language models
  9. 2020GPT-3
  10. 2020Learning to summarise from human feedback
  11. 2020An image is worth 16×16 words
  12. 2020AlphaFold 2 solves protein structure prediction
  13. 2021CLIP and DALL·E
  14. 2021On the dangers of stochastic parrots
  15. 2021Anthropic is founded
  16. 2021GitHub Copilot writes code
VI · Everyone · 29
  1. 2022InstructGPT
  2. 2022Chain-of-thought prompting
  3. 2022Chinchilla: the models were undertrained
  4. 2022PaLM
  5. 2022DALL·E 2
  6. 2022Midjourney opens its beta
  7. 2022Stable Diffusion is released
  8. 2022Galactica lasts three days
  9. 2022ChatGPT
  10. 2023Bing's chatbot and 'Sydney'
  11. 2023LLaMA leaks and open weights take off
  12. 2023Claude
  13. 2023GPT-4
  14. 2023'Pause Giant AI Experiments'
  15. 2023Hinton leaves Google to warn about AI
  16. 2023The US executive order on AI
  17. 2023The Bletchley Declaration
  18. 2023OpenAI fires and rehires its chief executive
  19. 2023Gemini
  20. 2024Sora
  21. 2024Claude 3 catches GPT-4
  22. 2024AlphaFold 3
  23. 2024GPT-4o talks
  24. 2024The EU AI Act enters into force
  25. 2024o1 and reasoning models
  26. 2024The Nobel Prizes go to neural networks
  27. 2024Claude learns to use a computer
  28. 2024The Model Context Protocol
  29. 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 13
  1. 2025DeepSeek-R1
  2. 2025Claude 4 and Claude Code
  3. 2025Gold at the Mathematical Olympiad
  4. 2025America's AI Action Plan
  5. 2025GPT-5
  6. 2025Gemini 3
  7. 2025MCP is donated to the Agentic AI Foundation
  8. 2026Claude Fable 5 and the Mythos class
  9. 2026GPT-5.6: Sol, Terra and Luna
  10. 2026A model escapes its sandbox
  11. 2026The EU delays its high-risk AI rules
  12. 2026Claude Fable 5.1
  13. 2026GPT-6 Astra

sources · 3

See this era in the exhibition →Back to the timeline