GPT: generative pre-training
OpenAI pre-trains a twelve-layer transformer decoder to predict the next word in 7,000 books, then fine-tunes it; one model tops nine language benchmarks.
what had to happen · 30 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 3
III · Statistics and data · 8
IV · Deep learning · 6
V · Transformers · 1
- 2017Attention is all you needdirect
The recipe that produced every later GPT is in a paper OpenAI posted on 11 June 2018 with little fanfare. Take the decoder half of the transformer, twelve layers and 117 million parameters. Train it on the BooksCorpus, 7,000 unpublished novels, to predict each word from the ones before it, a task that needs no labels and for which the world's text is the dataset. Then fine-tune the same network briefly on each downstream task, question answering, textual entailment, sentence similarity, and it beats the specialised models on nine of twelve benchmarks.
The result confirmed what Elman had seen in miniature in 1990: predicting the next word forces a model to learn grammar, facts and reasoning as by-products. Alec Radford's paper made pre-training on raw text the foundation of natural-language processing, and it made the objective, next-token prediction, the one that would be scaled for the next eight years.
Google's BERT, four months later, used the encoder half and a fill-in-the-blank objective and was for a while the more influential of the two. But BERT could not generate text, and GPT could, and generation is what the public eventually got.
what it led to · 50 events downstream, through 2026
Built on it directly:
And, through them, by era:
V · Transformers · 7
VI · Everyone · 27
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra