Learning to summarise from human feedback
OpenAI applies preference learning to GPT-style models: people pick the better of two summaries, a reward model learns the picks, and the language model is optimised against it.
what had to happen · 40 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 9
- 1948A mathematical theory of communication
- 1949Cells that fire together wire together
- 1950Programming a computer for playing chess
- 1958The perceptron learns
- 1959Samuel's checkers program coins 'machine learning'
- 1960ADALINE and the least-mean-squares rule
- 1965Moore's law
- 1969Perceptrons
- 1970Reverse-mode automatic differentiation
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 6
III · Statistics and data · 8
IV · Deep learning · 7
V · Transformers · 4
The 2017 preference-learning paper had trained simulated robots. This one, posted on 2 September 2020, trained a language model, and the recipe it settled on is the one still used. Take a pre-trained model and fine-tune it on human-written summaries of Reddit posts. Have people compare pairs of summaries and say which they prefer. Train a reward model to predict the preferences. Then optimise the language model, by reinforcement learning, to produce summaries the reward model scores highly, with a penalty for drifting too far from the original.
The summaries that resulted were preferred by people to the human-written references themselves, and the finding that a reward learned from a few thousand comparisons could outperform the standard objective was the practical result. The paper also documented what went wrong when the optimisation pushed too hard: the model found summaries that scored well and read badly, which its authors called over-optimisation and later work called reward hacking.
Fifteen months later the same team applied the procedure to instruction following, and InstructGPT was the outcome. The three-step recipe here, supervised fine-tuning, reward model, reinforcement learning, is what the acronym RLHF names.
what it led to · 31 events downstream, through 2026
Built on it directly:
- 2022InstructGPTVI
And, through them, by era:
VI · Everyone · 16
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Claude 3 catches GPT-4
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra