Skip to content
forward pass

o1 and reasoning models

OpenAI trains a model to think before it answers, spending more compute at inference on a hidden chain of thought; a second scaling axis opens and mathematics falls.

category
model
significance
4 of 5
people
Jakub Pachocki, Noam Brown, Ilya Sutskever
organisations
OpenAI

what had to happen · 53 events back to 1943

Every event this one built on, transitively, in order. The organism has the same path lit. Direct influences are marked.

00 · One neuron · 1

  1. 1943A logical calculus of nervous activity

I · Foundations · 11

  1. 1948A mathematical theory of communication
  2. 1949Cells that fire together wire together
  3. 1950Programming a computer for playing chess
  4. 1950Computing machinery and intelligence
  5. 1958The perceptron learns
  6. 1959Samuel's checkers program coins 'machine learning'
  7. 1960ADALINE and the least-mean-squares rule
  8. 1965Moore's law
  9. 1966ELIZA
  10. 1969Perceptrons
  11. 1970Reverse-mode automatic differentiation

W1 · The first winter · 2

  1. 1974Werbos applies backpropagation to neural networks
  2. 1980The Neocognitron

II · Connection · 3

  1. 1982The Hopfield network
  2. 1985The Boltzmann machine
  3. 1986Backpropagation

W2 · The second winter · 6

  1. 1988Temporal-difference learning
  2. 1989Q-learning
  3. 1989LeNet reads handwritten postcodes
  4. 1990Finding structure in time
  5. 1991The vanishing gradient problem
  6. 1992TD-Gammon reaches world-class backgammon

III · Statistics and data · 9

  1. 1997Long short-term memory
  2. 1998MNIST and LeNet-5
  3. 1999The first GPU
  4. 2003A neural probabilistic language model
  5. 2006Deep belief networks and the word 'deep'
  6. 2007CUDA
  7. 2009Deep learning moves to GPUs
  8. 2009ImageNet
  9. 2010Rectified linear units

IV · Deep learning · 9

  1. 2012Google Brain's network discovers cats
  2. 2012Dropout
  3. 2012AlexNet wins ImageNet
  4. 2013Deep Q-networks play Atari
  5. 2014Attention
  6. 2014Sequence to sequence learning
  7. 2015Batch normalisation
  8. 2015Residual networks
  9. 2016AlphaGo beats Lee Sedol

V · Transformers · 8

  1. 2017Attention is all you need
  2. 2017Deep reinforcement learning from human preferences
  3. 2017AlphaGo Zero learns from nothingdirect
  4. 2018GPT: generative pre-training
  5. 2019GPT-2 and the model too dangerous to release
  6. 2020Scaling laws for neural language models
  7. 2020GPT-3
  8. 2020Learning to summarise from human feedback

VI · Everyone · 4

  1. 2022InstructGPT
  2. 2022Chain-of-thought promptingdirect
  3. 2022ChatGPT
  4. 2023GPT-4direct

Chain-of-thought prompting had shown that models did better when they wrote out their reasoning. o1, previewed on 12 September 2024, was trained to do it. Using reinforcement learning on problems with checkable answers, mathematics, code, science, the model learned to produce long private chains of thought, to try approaches, notice errors and back up, before giving a reply. On a qualifying exam for the mathematical olympiad it solved 83 percent of problems where GPT-4o solved 13; on competitive programming it reached the 89th percentile of human entrants.

The company's chart showed accuracy rising smoothly with the compute spent at inference, on the same logarithmic axes as the 2020 scaling laws. Training compute had been the lever for a decade; now there was a second one, and a model could be made smarter by letting it think longer. The chain of thought was hidden from users, a choice OpenAI defended on safety and competitive grounds.

DeepSeek showed in January 2025 that the method could be reproduced cheaply and openly, and every laboratory shipped reasoning models within months. The gold-medal olympiad results of July 2025 came from their descendants.

what it led to · 8 events downstream, through 2026

Built on it directly:

  1. 2025DeepSeek-R1VII
  2. 2025Gold at the Mathematical OlympiadVII
  3. 2025GPT-5VII

And, through them, by era:

sources · 2

Trace the lineage of this event on the timeline →Back to the ledger