GPT-2 and the model too dangerous to release
OpenAI's 1.5-billion-parameter model writes coherent pages of text from a prompt; the lab withholds the full weights over misuse fears, and the argument about openness begins.
what had to happen · 31 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 3
III · Statistics and data · 8
IV · Deep learning · 6
V · Transformers · 2
- 2017Attention is all you need
- 2018GPT: generative pre-trainingdirect
GPT-2 was GPT scaled ten times, to 1.5 billion parameters, and trained on WebText, eight million pages linked from Reddit posts with at least three upvotes. Given a prompt it continued in the same style for paragraphs, inventing a news story about unicorns in the Andes that was quoted everywhere. It also answered questions, translated and summarised without being trained for any of those tasks, simply because the tasks appeared, in some form, in the text it had read. OpenAI called this zero-shot task transfer, and the finding that abilities emerged from scale alone set the agenda for the next five years.
The release, on 14 February 2019, was staged. Citing the risk of automated disinformation, OpenAI published only the smallest of four models and released the rest over nine months as it studied misuse. Critics called it a publicity stunt; supporters called it the first responsible-disclosure process for a model; both were partly right, and the full model, when it came, caused no visible harm.
GPT-2 also fixed the tokeniser. Its byte-pair encoding, which splits text into about 50,000 sub-word pieces, is what the tokens instrument on this site shows, and its successors in GPT-3, GPT-4 and beyond use the same scheme.
what it led to · 48 events downstream, through 2026
Built on it directly:
- 2020Scaling laws for neural language modelsV
- 2020GPT-3V
- 2020Learning to summarise from human feedbackV
And, through them, by era:
V · Transformers · 4
VI · Everyone · 27
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra