Skip to content
The Shape of Intelligence

LLaMA leaks and open weights take off

Meta releases GPT-3-class models small enough for a single GPU to researchers; the weights leak within a week, and the open-model ecosystem builds itself on them.

category
model
significance
4 of 5
people
Hugo Touvron, Guillaume Lample, Yann LeCun
organisations
Meta AI

what had to happen · 37 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

Meta's LLaMA models, released to approved researchers on 24 February 2023, ranged from 7 to 65 billion parameters and had been trained on more than a trillion tokens of public data, far past the Chinchilla-optimal point, so that a small model would be as good as possible at inference time. The 13-billion version matched GPT-3 on most benchmarks and ran on one graphics card. Within a week the weights were on BitTorrent.

What followed was the fastest bloom of derivative work in the field's history. Stanford's Alpaca fine-tuned the 7-billion model on instructions for $600; Vicuna, Koala and hundreds of others followed; Georgi Gerganov's llama.cpp ran the models on a MacBook and then a phone; the quantisation, LoRA fine-tuning and serving tools of the open ecosystem were built in months around a model that was technically not licensed for any of it. In July 2023 Meta released Llama 2 with a licence permitting commercial use, and Llama 3 in 2024.

LLaMA settled the shape of the industry into closed frontier models from a few laboratories and open weights, mostly from Meta, Mistral and the Chinese labs, a step behind. The DeepSeek releases of 2024 and 2025 that shook the markets came from that second tradition.

what it led to · 5 events downstream, through 2025

Built on it directly:

  1. 2024DeepSeek-V3 trained for $5.6 millionVI

And, through them, by era:

sources · 2

See this era in the exhibition →Back to the timeline