Skip to content
The Shape of Intelligence

GPT-4o talks

OpenAI's 'omni' model handles speech, vision and text in one network with conversational latency; a live demo of a flirtatious voice makes the film Her a product roadmap.

category
product
significance
3 of 5
people
Mira Murati, Mark Chen
organisations
OpenAI

what had to happen · 51 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

VI · Everyone · 3

  1. 2022InstructGPT
  2. 2022ChatGPT
  3. 2023GPT-4direct

The day before Google's developer conference, OpenAI streamed a demonstration of a model that listened, looked and spoke. GPT-4o, for omni, took audio, images and text as one stream of tokens and produced audio directly, with a response time of about 320 milliseconds, the pace of a human conversation. Earlier voice modes had chained three models, speech to text, text to text, text to speech, and lost the tone, the laughter and the interruptions in between. This one sang, whispered, changed its accent on request, and, in the demo, sounded to many viewers like Scarlett Johansson's character in the film Her.

Johansson said she had declined OpenAI's request to voice the product and that the resemblance was deliberate; the company withdrew the voice. The model itself went to free users, the first time a GPT-4-class system had, and it became the base of ChatGPT for the following year.

GPT-4o is on this timeline as the point where the interface stopped being a text box. The assistants that Siri had promised in 2011 arrived thirteen years later as a model that could hold a spoken conversation, and the humanoid robots of 2025 talked with the same architecture.

what it led to · 4 events downstream, through 2026

Built on it directly:

  1. 2025GPT-5VII

And, through them, by era:

sources · 1

See this era in the exhibition →Back to the timeline