Sequence to sequence learning
Sutskever, Vinyals and Le show that a large LSTM can translate English to French end to end, with no linguistic pipeline; text-in, text-out becomes the shape of the field.
what had to happen · 12 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 5
W1 · The first winter · 1
II · Connection · 1
- 1986Backpropagation
W2 · The second winter · 2
III · Statistics and data · 2
- 1997Long short-term memorydirect
- 2003A neural probabilistic language modeldirect
Machine translation in 2014 was a pipeline of statistical components, alignment, phrase tables, reordering, language models, engineered over twenty years. Ilya Sutskever, Oriol Vinyals and Quoc Le at Google replaced it with one network. An LSTM read the English sentence and produced a vector; a second LSTM read the vector and wrote the French, one word at a time. Trained on twelve million sentence pairs, four layers deep, with a trick of reversing the input so the beginnings of both sentences were close together, it matched the best phrase-based system on the WMT benchmark and beat it when the two were combined.
Kyunghyun Cho's group had published the same architecture in June; Bahdanau's attention had appeared nine days before Sutskever's preprint. Together the three papers established neural machine translation, and Google replaced its production system with one in 2016.
The larger point was the shape. A sequence goes in, a sequence comes out, and the network learns the mapping from examples: translation, summarisation, question answering, code generation and conversation are all the same problem. Every language model since is a sequence-to-sequence system, and Sutskever's insistence that scale would keep improving it took him to OpenAI as chief scientist the following year.
what it led to · 57 events downstream, through 2026
Built on it directly:
And, through them, by era:
V · Transformers · 12
- 2018GPT: generative pre-training
- 2018BERT
- 2019GPT-2 and the model too dangerous to release
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Learning to summarise from human feedback
- 2020An image is worth 16×16 words
- 2020AlphaFold 2 solves protein structure prediction
- 2021CLIP and DALL·E
- 2021On the dangers of stochastic parrots
- 2021Anthropic is founded
- 2021GitHub Copilot writes code
VI · Everyone · 29
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra