Skip to content
The Shape of Intelligence

· turning point

Scaling laws for neural language models

Kaplan and colleagues at OpenAI find that language-model loss falls as a smooth power law in parameters, data and compute across seven orders of magnitude; size becomes a plan.

category
theory
significance
5 of 5
people
Jared Kaplan, Sam McCandlish, Tom Henighan, Dario Amodei
organisations
OpenAI

what had to happen · 34 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

The paper posted on 23 January 2020 asked a plain empirical question: if you train transformers of different sizes on different amounts of text with different budgets of compute, how does the loss behave? The answer was a set of straight lines on log-log axes. Test loss fell as a power law in the number of parameters, in the size of the dataset and in the compute spent, each across many orders of magnitude, with no sign of levelling off. Architecture details mattered little; the exponents were what they were; and a model's performance could be predicted before it was trained.

That made scale a plan rather than a hope. Jared Kaplan's team, several of whom left to found Anthropic the next year, showed that for a fixed compute budget the best results came from a large model trained on relatively little data, stopped early, which is the recipe GPT-3 followed four months later. DeepMind's Chinchilla paper of 2022 corrected the exponents and found the data had been undervalued, but the shape of the result held.

The scaling laws are the reason the labs spent billions on compute, the reason the models kept improving, and the instrument on this page. Drag the compute and the loss follows the line.

what it led to · 46 events downstream, through 2026

Built on it directly:

  1. 2020GPT-3V
  2. 2021Anthropic is foundedV
  3. 2022Chinchilla: the models were undertrainedVI
  4. 2023GPT-4VI

And, through them, by era:

sources · 2

See this era in the exhibition →Back to the timeline