A neural probabilistic language model
Bengio's group learns a vector for every word and predicts the next word from the vectors of the last few; word embeddings and neural language models begin here.
what had to happen · 9 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 5
W1 · The first winter · 1
II · Connection · 1
- 1986Backpropagationdirect
W2 · The second winter · 1
- 1990Finding structure in timedirect
Language models in 2003 counted. An n-gram model estimated the probability of the next word from how often the preceding two or three words had been followed by it in a corpus, and it could not generalise: a sentence it had never seen was as unlikely as one that made no sense. Yoshua Bengio's group in Montréal proposed, in a paper first presented at NIPS in 2000 and published in full in 2003, to give every word a learned vector of a few dozen numbers and to predict the next word with a neural network from the vectors of the previous ones. Similar words would get similar vectors, and what the model learned about "cat" would transfer to "dog".
The paper introduced the word embedding, the idea that meaning could be a position in a space learned from prediction, and it beat the best n-gram models. It also took days to train on a few million words and was regarded as an elegant impracticality for most of a decade.
Word2vec in 2013 made the embeddings cheap and famous; the language models of the 2020s are this architecture with attention in place of the fixed window and a hundred billion parameters in place of a few hundred thousand. The task, predict the next word, has not changed since this paper.
what it led to · 60 events downstream, through 2026
Built on it directly:
- 2013Word2vecIV
- 2014AttentionIV
- 2014Sequence to sequence learningIV
- 2018GPT: generative pre-trainingV
And, through them, by era:
IV · Deep learning · 1
V · Transformers · 12
- 2017Attention is all you need
- 2018BERT
- 2019GPT-2 and the model too dangerous to release
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Learning to summarise from human feedback
- 2020An image is worth 16×16 words
- 2020AlphaFold 2 solves protein structure prediction
- 2021CLIP and DALL·E
- 2021On the dangers of stochastic parrots
- 2021Anthropic is founded
- 2021GitHub Copilot writes code
VI · Everyone · 29
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra