Skip to content
forward pass

InstructGPT

OpenAI fine-tunes GPT-3 with human feedback to follow instructions; a model a hundred times smaller is preferred by people to the original, and RLHF becomes the standard.

category
model
significance
4 of 5
people
Long Ouyang, Jeff Wu, Jan Leike, Paul Christiano
organisations
OpenAI

what had to happen · 45 events back to 1943

Every event this one built on, transitively, in order. The organism has the same path lit. Direct influences are marked.

00 · One neuron · 1

  1. 1943A logical calculus of nervous activity

I · Foundations · 9

  1. 1948A mathematical theory of communication
  2. 1949Cells that fire together wire together
  3. 1950Programming a computer for playing chess
  4. 1958The perceptron learns
  5. 1959Samuel's checkers program coins 'machine learning'
  6. 1960ADALINE and the least-mean-squares rule
  7. 1965Moore's law
  8. 1969Perceptrons
  9. 1970Reverse-mode automatic differentiation

W1 · The first winter · 2

  1. 1974Werbos applies backpropagation to neural networks
  2. 1980The Neocognitron

II · Connection · 3

  1. 1982The Hopfield network
  2. 1985The Boltzmann machine
  3. 1986Backpropagation

W2 · The second winter · 6

  1. 1988Temporal-difference learning
  2. 1989Q-learning
  3. 1989LeNet reads handwritten postcodes
  4. 1990Finding structure in time
  5. 1991The vanishing gradient problem
  6. 1992TD-Gammon reaches world-class backgammon

III · Statistics and data · 9

  1. 1997Long short-term memory
  2. 1998MNIST and LeNet-5
  3. 1999The first GPU
  4. 2003A neural probabilistic language model
  5. 2006Deep belief networks and the word 'deep'
  6. 2007CUDA
  7. 2009Deep learning moves to GPUs
  8. 2009ImageNet
  9. 2010Rectified linear units

IV · Deep learning · 8

  1. 2012Google Brain's network discovers cats
  2. 2012Dropout
  3. 2012AlexNet wins ImageNet
  4. 2013Deep Q-networks play Atari
  5. 2014Attention
  6. 2014Sequence to sequence learning
  7. 2015Batch normalisation
  8. 2015Residual networks

V · Transformers · 7

  1. 2017Attention is all you need
  2. 2017Deep reinforcement learning from human preferencesdirect
  3. 2018GPT: generative pre-training
  4. 2019GPT-2 and the model too dangerous to release
  5. 2020Scaling laws for neural language models
  6. 2020GPT-3direct
  7. 2020Learning to summarise from human feedbackdirect

GPT-3 continued text. Asked to explain the moon landing to a six-year-old, it might produce a list of other things to explain to six-year-olds, because that is a plausible continuation. InstructGPT, announced on 27 January 2022, was the same model taught to do what it was asked. Forty contractors wrote demonstrations of good responses to prompts from the API; the model was fine-tuned on them; the contractors then ranked the model's outputs, a reward model learned the rankings, and the model was optimised against the reward with a leash to keep it from drifting.

Labellers preferred the responses of the 1.3-billion-parameter InstructGPT to those of the 175-billion-parameter GPT-3. The model made up facts less often, was less toxic, and followed instructions it had never seen in training, in languages that were barely present in the feedback. The cost of the alignment step was a tiny fraction of the cost of pre-training.

InstructGPT became the default model of the API in January 2022, and a sibling trained the same way on conversation became ChatGPT ten months later. The recipe has been used on every chat model since. The instrument on this page lets you be the labeller.

what it led to · 30 events downstream, through 2026

Built on it directly:

  1. 2022ChatGPTVI
  2. 2023ClaudeVI
  3. 2023GPT-4VI

And, through them, by era:

sources · 2

Trace the lineage of this event on the timeline →Back to the ledger