· turning point
Scaling laws for neural language models
Kaplan and colleagues at OpenAI find that language-model loss falls as a smooth power law in parameters, data and compute across seven orders of magnitude; size becomes a plan.
what had to happen · 34 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 3
III · Statistics and data · 9
IV · Deep learning · 7
V · Transformers · 3
The paper posted on 23 January 2020 asked a plain empirical question: if you train transformers of different sizes on different amounts of text with different budgets of compute, how does the loss behave? The answer was a set of straight lines on log-log axes. Test loss fell as a power law in the number of parameters, in the size of the dataset and in the compute spent, each across many orders of magnitude, with no sign of levelling off. Architecture details mattered little; the exponents were what they were; and a model's performance could be predicted before it was trained.
That made scale a plan rather than a hope. Jared Kaplan's team, several of whom left to found Anthropic the next year, showed that for a fixed compute budget the best results came from a large model trained on relatively little data, stopped early, which is the recipe GPT-3 followed four months later. DeepMind's Chinchilla paper of 2022 corrected the exponents and found the data had been undervalued, but the shape of the result held.
The scaling laws are the reason the labs spent billions on compute, the reason the models kept improving, and the instrument on this page. Drag the compute and the loss follows the line.
what it led to · 46 events downstream, through 2026
Built on it directly:
- 2020GPT-3V
- 2021Anthropic is foundedV
- 2022Chinchilla: the models were undertrainedVI
- 2023GPT-4VI
And, through them, by era:
V · Transformers · 3
VI · Everyone · 25
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra