Sora
OpenAI shows minute-long videos generated from text by a diffusion transformer over spacetime patches; film and advertising begin to plan around it.
what had to happen · 41 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 3
III · Statistics and data · 9
IV · Deep learning · 8
V · Transformers · 8
VI · Everyone · 1
- 2022Stable Diffusion is releaseddirect
On 15 February 2024 OpenAI published a page of videos: a woman walking down a Tokyo street at night, woolly mammoths in snow, a drone shot of a coastline, each up to a minute long, at 1080p, with consistent objects and plausible physics, and each made from a paragraph of text. Sora was a diffusion model whose network was a transformer rather than a U-Net, working on patches of video compressed in space and time, the vision transformer's idea carried into the third and fourth dimensions.
The demonstration was not a product; access went to a few artists and red-teamers, and the model was not generally released until December. Its effect was immediate anyway. Tyler Perry paused an $800 million studio expansion; stock-footage companies and animators reconsidered their businesses; and the video models that Google, Runway, Kling and others released over the following eighteen months were measured against a page of clips most people could not use.
Sora 2, released with a social app on 30 September 2025, added sound and let users insert themselves into scenes. Text-to-video is on this timeline as the point at which generative models reached the last medium, and at which the question of what a recording is evidence of became unanswerable.
what it led to · 0 events downstream
A leaf, for now. Nothing in the archive has built on it yet.