Skip to content
forward pass

GPT-4o talks

OpenAI's 'omni' model handles speech, vision and text in one network with conversational latency; a live demo of a flirtatious voice makes the film Her a product roadmap.

category
product
significance
3 of 5
people
Mira Murati, Mark Chen
organisations
OpenAI

what had to happen · 51 events back to 1943

Every event this one built on, transitively, in order. The organism has the same path lit. Direct influences are marked.

00 · One neuron · 1

  1. 1943A logical calculus of nervous activity

I · Foundations · 11

  1. 1948A mathematical theory of communication
  2. 1949Cells that fire together wire together
  3. 1950Programming a computer for playing chess
  4. 1950Computing machinery and intelligence
  5. 1958The perceptron learns
  6. 1959Samuel's checkers program coins 'machine learning'
  7. 1960ADALINE and the least-mean-squares rule
  8. 1965Moore's law
  9. 1966ELIZA
  10. 1969Perceptrons
  11. 1970Reverse-mode automatic differentiation

W1 · The first winter · 2

  1. 1974Werbos applies backpropagation to neural networks
  2. 1980The Neocognitron

II · Connection · 3

  1. 1982The Hopfield network
  2. 1985The Boltzmann machine
  3. 1986Backpropagation

W2 · The second winter · 6

  1. 1988Temporal-difference learning
  2. 1989Q-learning
  3. 1989LeNet reads handwritten postcodes
  4. 1990Finding structure in time
  5. 1991The vanishing gradient problem
  6. 1992TD-Gammon reaches world-class backgammon

III · Statistics and data · 9

  1. 1997Long short-term memory
  2. 1998MNIST and LeNet-5
  3. 1999The first GPU
  4. 2003A neural probabilistic language model
  5. 2006Deep belief networks and the word 'deep'
  6. 2007CUDA
  7. 2009Deep learning moves to GPUs
  8. 2009ImageNet
  9. 2010Rectified linear units

IV · Deep learning · 9

  1. 2012Google Brain's network discovers cats
  2. 2012Dropout
  3. 2012AlexNet wins ImageNet
  4. 2013Deep Q-networks play Atari
  5. 2014Attention
  6. 2014Sequence to sequence learning
  7. 2015Batch normalisation
  8. 2015Residual networks
  9. 2016WaveNetdirect

V · Transformers · 7

  1. 2017Attention is all you need
  2. 2017Deep reinforcement learning from human preferences
  3. 2018GPT: generative pre-training
  4. 2019GPT-2 and the model too dangerous to release
  5. 2020Scaling laws for neural language models
  6. 2020GPT-3
  7. 2020Learning to summarise from human feedback

VI · Everyone · 3

  1. 2022InstructGPT
  2. 2022ChatGPT
  3. 2023GPT-4direct

The day before Google's developer conference, OpenAI streamed a demonstration of a model that listened, looked and spoke. GPT-4o, for omni, took audio, images and text as one stream of tokens and produced audio directly, with a response time of about 320 milliseconds, the pace of a human conversation. Earlier voice modes had chained three models, speech to text, text to text, text to speech, and lost the tone, the laughter and the interruptions in between. This one sang, whispered, changed its accent on request, and, in the demo, sounded to many viewers like Scarlett Johansson's character in the film Her.

Johansson said she had declined OpenAI's request to voice the product and that the resemblance was deliberate; the company withdrew the voice. The model itself went to free users, the first time a GPT-4-class system had, and it became the base of ChatGPT for the following year.

GPT-4o is on this timeline as the point where the interface stopped being a text box. The assistants that Siri had promised in 2011 arrived thirteen years later as a model that could hold a spoken conversation, and the humanoid robots of 2025 talked with the same architecture.

what it led to · 4 events downstream, through 2026

Built on it directly:

  1. 2025GPT-5VII

And, through them, by era:

sources · 1

Trace the lineage of this event on the timeline →Back to the ledger