Skip to content
The Shape of Intelligence

Google reveals the TPU

Google discloses that a custom chip for neural-network inference has been running in its data centres for a year; the hardware race for AI moves beyond GPUs.

category
hardware
significance
3 of 5
people
Norman Jouppi, Jeff Dean
organisations
Google

what had to happen · 23 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

At its I/O conference on 18 May 2016, Google mentioned that its data centres had for more than a year been running a chip it had designed itself, the Tensor Processing Unit, built to do one thing: the eight-bit matrix multiplications of neural-network inference. It had powered the search ranking model RankBrain and, in March, AlphaGo's match against Lee Sedol. The 2017 paper by Norman Jouppi's team put the numbers on it, fifteen to thirty times the performance of a contemporary GPU or CPU on Google's workloads, at thirty to eighty times the performance per watt.

The disclosure changed the hardware market. Until then, NVIDIA's graphics processors, designed for games and adapted for learning, were the only serious option. The TPU showed that the arithmetic of deep learning was regular enough to deserve its own silicon, and that a company running models at Google's scale would save enough power to justify building it. Amazon, Microsoft, Meta, Tesla and a dozen startups followed with chips of their own.

The second-generation TPU of 2017 could train as well as infer, and the pods of thousands of them trained BERT, the transformer's successors and Gemini. The largest language models in the world have been trained about equally on Google's chips and NVIDIA's.

what it led to · 3 events downstream, through 2025

Built on it directly:

  1. 2022PaLMVI

And, through them, by era:

VI · Everyone · 1
  1. 2023Gemini
VII · Agents · 1
  1. 2025Gemini 3

sources · 2

See this era in the exhibition →Back to the timeline