Skip to content
The Shape of Intelligence

Deep learning moves to GPUs

Raina, Madhavan and Ng train deep belief networks on graphics cards seventy times faster than on CPUs; the hardware and the method find each other.

category
theory
significance
3 of 5
people
Rajat Raina, Anand Madhavan, Andrew Ng
organisations
Stanford University

what had to happen · 14 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

Andrew Ng's group at Stanford wanted to train unsupervised models with a hundred million parameters, and on the CPUs of 2008 that took weeks. Their paper at ICML in June 2009 reported what happened when they wrote the training code in CUDA for an NVIDIA GTX 280. Deep belief networks trained up to seventy times faster; sparse coding, fifteen times. Models that had been out of reach were trained in a day.

The arithmetic was not surprising to anyone who had looked. A graphics card of the time had 240 cores and a memory bandwidth ten times that of a CPU, and neural-network training is almost entirely dense linear algebra. What the paper did was demonstrate it on the models people were actually trying to build, and publish the speed-ups in a venue the field read.

Within three years the practice was universal. Dan Cireşan's group in Switzerland set records on MNIST and traffic signs with GPU-trained networks in 2010 and 2011; Alex Krizhevsky trained AlexNet on two consumer cards in 2012. The lesson Ng drew, that scale in compute and data mattered more than cleverness in models, took him to Google Brain in 2011 and its billion-parameter experiments.

what it led to · 49 events downstream, through 2026

Built on it directly:

  1. 2012Google Brain's network discovers catsIV

And, through them, by era:

sources · 1

See this era in the exhibition →Back to the timeline