Google reveals the TPU
Google discloses that a custom chip for neural-network inference has been running in its data centres for a year; the hardware race for AI moves beyond GPUs.
what had to happen · 23 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 3
III · Statistics and data · 6
- 1998MNIST and LeNet-5
- 1999The first GPU
- 2006Deep belief networks and the word 'deep'
- 2007CUDAdirect
- 2009ImageNet
- 2010Rectified linear units
IV · Deep learning · 2
- 2012Dropout
- 2012AlexNet wins ImageNetdirect
At its I/O conference on 18 May 2016, Google mentioned that its data centres had for more than a year been running a chip it had designed itself, the Tensor Processing Unit, built to do one thing: the eight-bit matrix multiplications of neural-network inference. It had powered the search ranking model RankBrain and, in March, AlphaGo's match against Lee Sedol. The 2017 paper by Norman Jouppi's team put the numbers on it, fifteen to thirty times the performance of a contemporary GPU or CPU on Google's workloads, at thirty to eighty times the performance per watt.
The disclosure changed the hardware market. Until then, NVIDIA's graphics processors, designed for games and adapted for learning, were the only serious option. The TPU showed that the arithmetic of deep learning was regular enough to deserve its own silicon, and that a company running models at Google's scale would save enough power to justify building it. Amazon, Microsoft, Meta, Tesla and a dozen startups followed with chips of their own.
The second-generation TPU of 2017 could train as well as infer, and the pods of thousands of them trained BERT, the transformer's successors and Gemini. The largest language models in the world have been trained about equally on Google's chips and NVIDIA's.
what it led to · 3 events downstream, through 2025
Built on it directly:
- 2022PaLMVI
And, through them, by era: