Google Brain's network discovers cats
A billion-parameter network trained on ten million YouTube frames across 16,000 cores learns, unsupervised, a neuron that fires for cat faces; scale enters the vocabulary.
what had to happen · 15 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 1
II · Connection · 3
III · Statistics and data · 4
- 1999The first GPU
- 2006Deep belief networks and the word 'deep'direct
- 2007CUDA
- 2009Deep learning moves to GPUsdirect
Google Brain began in 2011 as a collaboration between Andrew Ng, Jeff Dean and Greg Corrado to find out what happened when the deep networks of Ng's Stanford group were trained on Google's computers. The answer, presented at ICML in June 2012 and on the front page of the New York Times, was a network with a billion connections, trained for three days on 16,000 processor cores using frames from ten million YouTube videos, with no labels at all. Among its top-level units was one that responded to cat faces, and another to human faces, which nobody had asked for.
The result was modest as vision, the network's accuracy on ImageNet categories was 15.8 percent, and enormous as a demonstration. It showed that a large enough network with enough data would organise its own concepts, and that the limiting factor was compute. Within months Google had built the infrastructure, DistBelief, that became TensorFlow, and had hired Geoffrey Hinton.
The paper also fixed a number in the public mind. A billion parameters, in 2012, was a headline; GPT-3 would have 175 billion eight years later and the frontier models of the 2020s more than a trillion. The cat neuron was the first widely reported evidence that the way to make networks smarter was to make them bigger.
what it led to · 48 events downstream, through 2026
Built on it directly:
And, through them, by era:
V · Transformers · 5
VI · Everyone · 27
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra