Skip to content
The Shape of Intelligence

PaLM

Google trains a 540-billion-parameter model across two TPU pods and reports emergent abilities that appear only at scale, explaining jokes and reasoning through problems.

category
model
significance
3 of 5
people
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin
organisations
Google Research

what had to happen · 38 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

The Pathways Language Model, announced on 4 April 2022, was the largest dense transformer disclosed at the time, 540 billion parameters trained on 780 billion tokens across 6,144 TPU v4 chips in two pods, using a new system, Pathways, to spread one model across data centres. Its paper is 87 pages, and much of it is a catalogue of things the model could do that its smaller versions could not: explain why a joke was funny, follow chains of reasoning through multi-step problems, and, with chain-of-thought prompting, match fine-tuned models on maths.

The paper made "emergent abilities" a term of art, with a graph of tasks whose accuracy sat at chance until a certain scale and then jumped. Whether the jumps were real or artefacts of how the tasks were scored became a debate that ran for two years. Either way, PaLM was Google's demonstration that it could match OpenAI's scale, eight months before ChatGPT made the question commercial.

PaLM 2 in May 2023 ran Bard; Gemini replaced both in December. The model is on this timeline as the high point of the pure-scale era, before Chinchilla's data correction and human feedback changed what the laboratories optimised.

what it led to · 2 events downstream, through 2025

Built on it directly:

  1. 2023GeminiVI

And, through them, by era:

VII · Agents · 1
  1. 2025Gemini 3

sources · 1

See this era in the exhibition →Back to the timeline