Skip to content
The Shape of Intelligence

DALL·E 2

OpenAI's second image model combines CLIP with diffusion to produce photorealistic pictures from text; the astronaut on a horse goes everywhere and a waiting list forms.

category
model
significance
3 of 5
people
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol
organisations
OpenAI

what had to happen · 40 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

DALL·E 2, announced on 6 April 2022, was the moment machine-made images became good enough to be shared for their own sake. Its architecture was in two parts. A prior turned a text description into a CLIP image embedding; a diffusion decoder turned the embedding into a picture at 1,024 pixels square, in any style requested. The example that went round the world was an astronaut riding a horse in a photorealistic style, and the model produced a plausible one on the first try.

Access was by invitation for months, and the waiting list, the watermark in the corner, and the rules about faces and violence were a rehearsal for the debates of the next two years about who could make what. Artists objected to the training data; stock-photo companies banned the outputs; the model would not draw public figures.

Google's Imagen the following month and Midjourney's open beta in July matched it, and Stable Diffusion in August gave the capability away. DALL·E 2 was the first, and the one that fixed in the public mind what a text-to-image system was.

what it led to · 0 events downstream

A leaf, for now. Nothing in the archive has built on it yet.

sources · 1

See this era in the exhibition →Back to the timeline