DALL·E 2
OpenAI's second image model combines CLIP with diffusion to produce photorealistic pictures from text; the astronaut on a horse goes everywhere and a waiting list forms.
what had to happen · 40 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 3
III · Statistics and data · 9
IV · Deep learning · 8
V · Transformers · 8
DALL·E 2, announced on 6 April 2022, was the moment machine-made images became good enough to be shared for their own sake. Its architecture was in two parts. A prior turned a text description into a CLIP image embedding; a diffusion decoder turned the embedding into a picture at 1,024 pixels square, in any style requested. The example that went round the world was an astronaut riding a horse in a photorealistic style, and the model produced a plausible one on the first try.
Access was by invitation for months, and the waiting list, the watermark in the corner, and the rules about faces and violence were a rehearsal for the debates of the next two years about who could make what. Artists objected to the training data; stock-photo companies banned the outputs; the model would not draw public figures.
Google's Imagen the following month and Midjourney's open beta in July matched it, and Stable Diffusion in August gave the capability away. DALL·E 2 was the first, and the one that fixed in the public mind what a text-to-image system was.
what it led to · 0 events downstream
A leaf, for now. Nothing in the archive has built on it yet.