Skip to content
The Shape of Intelligence

Sora

OpenAI shows minute-long videos generated from text by a diffusion transformer over spacetime patches; film and advertising begin to plan around it.

category
model
significance
3 of 5
people
Tim Brooks, Bill Peebles
organisations
OpenAI

what had to happen · 41 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

On 15 February 2024 OpenAI published a page of videos: a woman walking down a Tokyo street at night, woolly mammoths in snow, a drone shot of a coastline, each up to a minute long, at 1080p, with consistent objects and plausible physics, and each made from a paragraph of text. Sora was a diffusion model whose network was a transformer rather than a U-Net, working on patches of video compressed in space and time, the vision transformer's idea carried into the third and fourth dimensions.

The demonstration was not a product; access went to a few artists and red-teamers, and the model was not generally released until December. Its effect was immediate anyway. Tyler Perry paused an $800 million studio expansion; stock-footage companies and animators reconsidered their businesses; and the video models that Google, Runway, Kling and others released over the following eighteen months were measured against a page of clips most people could not use.

Sora 2, released with a social app on 30 September 2025, added sound and let users insert themselves into scenes. Text-to-video is on this timeline as the point at which generative models reached the last medium, and at which the question of what a recording is evidence of became unanswerable.

what it led to · 0 events downstream

A leaf, for now. Nothing in the archive has built on it yet.

sources · 1

See this era in the exhibition →Back to the timeline