PaLM
Google trains a 540-billion-parameter model across two TPU pods and reports emergent abilities that appear only at scale, explaining jokes and reasoning through problems.
what had to happen · 38 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 3
III · Statistics and data · 9
IV · Deep learning · 8
V · Transformers · 5
VI · Everyone · 1
- 2022Chain-of-thought promptingdirect
The Pathways Language Model, announced on 4 April 2022, was the largest dense transformer disclosed at the time, 540 billion parameters trained on 780 billion tokens across 6,144 TPU v4 chips in two pods, using a new system, Pathways, to spread one model across data centres. Its paper is 87 pages, and much of it is a catalogue of things the model could do that its smaller versions could not: explain why a joke was funny, follow chains of reasoning through multi-step problems, and, with chain-of-thought prompting, match fine-tuned models on maths.
The paper made "emergent abilities" a term of art, with a graph of tasks whose accuracy sat at chance until a certain scale and then jumped. Whether the jumps were real or artefacts of how the tasks were scored became a debate that ran for two years. Either way, PaLM was Google's demonstration that it could match OpenAI's scale, eight months before ChatGPT made the question commercial.
PaLM 2 in May 2023 ran Bard; Gemini replaced both in December. The model is on this timeline as the high point of the pure-scale era, before Chinchilla's data correction and human feedback changed what the laboratories optimised.
what it led to · 2 events downstream, through 2025
Built on it directly:
- 2023GeminiVI
And, through them, by era:
VII · Agents · 1
- 2025Gemini 3