Skip to content
forward pass

LLaMA leaks and open weights take off

Meta releases GPT-3-class models small enough for a single GPU to researchers; the weights leak within a week, and the open-model ecosystem builds itself on them.

category
model
significance
4 of 5
people
Hugo Touvron, Guillaume Lample, Yann LeCun
organisations
Meta AI

what had to happen · 37 events back to 1943

Every event this one built on, transitively, in order. The organism has the same path lit. Direct influences are marked.

Meta's LLaMA models, released to approved researchers on 24 February 2023, ranged from 7 to 65 billion parameters and had been trained on more than a trillion tokens of public data, far past the Chinchilla-optimal point, so that a small model would be as good as possible at inference time. The 13-billion version matched GPT-3 on most benchmarks and ran on one graphics card. Within a week the weights were on BitTorrent.

What followed was the fastest bloom of derivative work in the field's history. Stanford's Alpaca fine-tuned the 7-billion model on instructions for $600; Vicuna, Koala and hundreds of others followed; Georgi Gerganov's llama.cpp ran the models on a MacBook and then a phone; the quantisation, LoRA fine-tuning and serving tools of the open ecosystem were built in months around a model that was technically not licensed for any of it. In July 2023 Meta released Llama 2 with a licence permitting commercial use, and Llama 3 in 2024.

LLaMA settled the shape of the industry into closed frontier models from a few laboratories and open weights, mostly from Meta, Mistral and the Chinese labs, a step behind. The DeepSeek releases of 2024 and 2025 that shook the markets came from that second tradition.

what it led to · 5 events downstream, through 2025

Built on it directly:

  1. 2024DeepSeek-V3 trained for $5.6 millionVI

And, through them, by era:

sources · 2

Trace the lineage of this event on the timeline →Back to the ledger