Deep reinforcement learning from human preferences
Christiano and colleagues train agents from a human's choices between pairs of video clips instead of a coded reward; the method that will align chatbots is born.
what had to happen · 30 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 9
- 1948A mathematical theory of communication
- 1949Cells that fire together wire together
- 1950Programming a computer for playing chess
- 1958The perceptron learns
- 1959Samuel's checkers program coins 'machine learning'
- 1960ADALINE and the least-mean-squares rule
- 1965Moore's law
- 1969Perceptrons
- 1970Reverse-mode automatic differentiation
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 6
III · Statistics and data · 6
IV · Deep learning · 3
- 2012Dropout
- 2012AlexNet wins ImageNet
- 2013Deep Q-networks play Ataridirect
Reinforcement learning needs a reward, and for most things people want, no one can write the reward down. Paul Christiano's paper, a collaboration between OpenAI and DeepMind posted on 12 June 2017, the same day as the transformer, proposed getting it from people a different way. Show a human two short clips of an agent's behaviour and ask which is better. Train a reward model to predict those judgements. Train the agent against the reward model. Repeat, asking the human about the cases the reward model is least sure of. With about an hour of a person's time, a simulated robot learned to do a backflip, a behaviour that nobody had managed to specify as a formula.
The paper was a safety paper. Its authors, several of whom later founded Anthropic, were looking for a way to make systems do what people meant rather than what they had literally been told, the problem Bostrom had made famous.
Three years later the same recipe was applied to language: a model writes two summaries, a person picks the better one, and the model is trained towards the preference. That became reinforcement learning from human feedback, the step that turned GPT-3 into InstructGPT and InstructGPT into ChatGPT. The chatbot that talks to you politely was tuned by the method in this paper.
what it led to · 32 events downstream, through 2026
Built on it directly:
And, through them, by era:
VI · Everyone · 16
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Claude 3 catches GPT-4
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
VII · Agents · 13
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra