Denoising diffusion probabilistic models
Ho, Jain and Abbeel simplify the 2015 diffusion recipe into predicting the noise, and match adversarial networks on image quality; the generative field changes course.
what had to happen · 5 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 1
II · Connection · 2
IV · Deep learning · 1
- 2015Diffusion modelsdirect
Sohl-Dickstein's diffusion models of 2015 had been correct and impractical. Jonathan Ho's paper, posted on 19 June 2020, made two changes. It trained the network on a much simpler target, the noise that had been added at each step rather than the full reverse distribution, which turned out to be equivalent to a weighted form of the original objective and far easier to learn. And it used a large U-Net, the image-to-image architecture from medical segmentation, with attention inside it. On the CIFAR-10 benchmark the results matched the best adversarial networks, and on faces they were, to most eyes, better.
Diffusion had two properties GANs lacked. Training was stable, a single loss going down, with none of the collapse and oscillation of the adversarial game. And the model covered the whole distribution rather than the parts it found easiest to fake. Within a year, Prafulla Dhariwal and Alex Nichol had shown diffusion beating GANs on ImageNet, and OpenAI's GLIDE and DALL-E 2, Google's Imagen and Stability's Stable Diffusion were all built on it.
The film of noise resolving into a picture, which is how everyone now imagines a machine making an image, is this algorithm. The instrument on this page shows the forward process on a real image and the reverse as an illustration.
what it led to · 5 events downstream, through 2024
Built on it directly:
- 2022DALL·E 2VI
- 2022Midjourney opens its betaVI
- 2022Stable Diffusion is releasedVI
- 2024AlphaFold 3VI
And, through them, by era:
VI · Everyone · 1
- 2024Sora