Skip to content
forward pass

Claude 3 catches GPT-4

Anthropic's Haiku, Sonnet and Opus models arrive with vision and a 200,000-token window; Opus tops the leaderboards, and for the first time OpenAI is not alone at the front.

category
model
significance
3 of 5
people
Dario Amodei, Jared Kaplan
organisations
Anthropic

what had to happen · 52 events back to 1943

Every event this one built on, transitively, in order. The organism has the same path lit. Direct influences are marked.

00 · One neuron · 1

  1. 1943A logical calculus of nervous activity

I · Foundations · 9

  1. 1948A mathematical theory of communication
  2. 1949Cells that fire together wire together
  3. 1950Programming a computer for playing chess
  4. 1958The perceptron learns
  5. 1959Samuel's checkers program coins 'machine learning'
  6. 1960ADALINE and the least-mean-squares rule
  7. 1965Moore's law
  8. 1969Perceptrons
  9. 1970Reverse-mode automatic differentiation

W1 · The first winter · 2

  1. 1974Werbos applies backpropagation to neural networks
  2. 1980The Neocognitron

II · Connection · 3

  1. 1982The Hopfield network
  2. 1985The Boltzmann machine
  3. 1986Backpropagation

W2 · The second winter · 6

  1. 1988Temporal-difference learning
  2. 1989Q-learning
  3. 1989LeNet reads handwritten postcodes
  4. 1990Finding structure in time
  5. 1991The vanishing gradient problem
  6. 1992TD-Gammon reaches world-class backgammon

III · Statistics and data · 10

  1. 1997Long short-term memory
  2. 1998MNIST and LeNet-5
  3. 1999The first GPU
  4. 2003A neural probabilistic language model
  5. 2006Deep belief networks and the word 'deep'
  6. 2007CUDA
  7. 2009Deep learning moves to GPUs
  8. 2009ImageNet
  9. 2010Rectified linear units
  10. 2010DeepMind is founded

IV · Deep learning · 11

  1. 2012Google Brain's network discovers cats
  2. 2012Dropout
  3. 2012AlexNet wins ImageNet
  4. 2013Deep Q-networks play Atari
  5. 2014Google buys DeepMind
  6. 2014Superintelligence
  7. 2014Attention
  8. 2014Sequence to sequence learning
  9. 2015Batch normalisation
  10. 2015Residual networks
  11. 2015OpenAI is founded

V · Transformers · 8

  1. 2017Attention is all you need
  2. 2017Deep reinforcement learning from human preferences
  3. 2018GPT: generative pre-training
  4. 2019GPT-2 and the model too dangerous to release
  5. 2020Scaling laws for neural language models
  6. 2020GPT-3
  7. 2020Learning to summarise from human feedback
  8. 2021Anthropic is founded

VI · Everyone · 2

  1. 2022InstructGPT
  2. 2023Claudedirect

The Claude 3 family, released on 4 March 2024, came in three sizes named for lengths of poetry, Haiku, Sonnet and Opus, and the largest was the first model to beat GPT-4 across the standard benchmarks in the year since GPT-4's release. All three could read images and documents, held 200,000 tokens of context, and were markedly less inclined than their predecessor to refuse harmless requests, a complaint that had defined Claude 2.

Two things about the release stuck. In a needle-in-a-haystack test, in which a fact is hidden in a long document, Opus not only found the fact but remarked that it appeared to have been inserted as a test, which was reported as self-awareness and was, more prosaically, evidence of how much the models had learned about the tests they were given. And Anthropic's model card spent pages on the model's own reports of its experience, treating the question as open.

Claude 3.5 Sonnet in June 2024 became, by most measures, the best model for writing code, and the family's later versions, 3.7 and 4, ran the coding agents of 2025. The March 2024 release is the point at which the frontier had three laboratories abreast rather than one ahead.

what it led to · 6 events downstream, through 2026

Built on it directly:

  1. 2024Claude learns to use a computerVI
  2. 2025Claude 4 and Claude CodeVII

And, through them, by era:

sources · 1

Trace the lineage of this event on the timeline →Back to the ledger