· turning point
Attention is all you need
Eight Google researchers drop recurrence entirely and build a sequence model from attention alone; the transformer trains in parallel, scales without limit, and becomes the architecture of everything.
what had to happen · 29 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 6
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 3
III · Statistics and data · 8
IV · Deep learning · 6
- 2012Dropout
- 2012AlexNet wins ImageNet
- 2014Attentiondirect
- 2014Sequence to sequence learningdirect
- 2015Batch normalisation
- 2015Residual networksdirect
The paper posted on 12 June 2017 was a translation paper. Its eight authors, listed in an order they said was random, had found that the recurrent networks everyone used for sequences were unnecessary. Bahdanau's attention, which let a decoder look back over the source, could be turned inward: every position in a sentence attends to every other position, computing what to weigh from learned queries and keys, and the whole layer runs at once rather than word by word. Stack six of those layers, with residual connections and normalisation between them, add a position signal so the model knows word order, and you have the transformer. It beat the best translation systems on English–German and English–French and trained in a fraction of the time.
Speed was the point. A recurrent network processes a thousand-word document in a thousand sequential steps; a transformer processes it in one, which means it can use every core of a GPU pod, which means it can be trained on far more data. Scale, which the field had been suspecting was the answer since 2012, became possible for language.
Every model on this timeline after 2017 that reads or writes is a transformer: BERT, the GPTs, AlphaFold 2, Stable Diffusion's text encoder, the vision transformer, the models that drive cars and fold proteins and write code. The paper has been cited more than 150,000 times. The instrument below runs one on the sentence you type and shows you the attention.
what it led to · 55 events downstream, through 2026
Built on it directly:
- 2018GPT: generative pre-trainingV
- 2018BERTV
- 2020An image is worth 16×16 wordsV
- 2020AlphaFold 2 solves protein structure predictionV
And, through them, by era:
V · Transformers · 8
VI · Everyone · 29
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra