Skip to content
The Shape of Intelligence

Deep reinforcement learning from human preferences

Christiano and colleagues train agents from a human's choices between pairs of video clips instead of a coded reward; the method that will align chatbots is born.

category
theory
significance
4 of 5
people
Paul Christiano, Jan Leike, Tom Brown, Dario Amodei
organisations
OpenAI, DeepMind

what had to happen · 30 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

Reinforcement learning needs a reward, and for most things people want, no one can write the reward down. Paul Christiano's paper, a collaboration between OpenAI and DeepMind posted on 12 June 2017, the same day as the transformer, proposed getting it from people a different way. Show a human two short clips of an agent's behaviour and ask which is better. Train a reward model to predict those judgements. Train the agent against the reward model. Repeat, asking the human about the cases the reward model is least sure of. With about an hour of a person's time, a simulated robot learned to do a backflip, a behaviour that nobody had managed to specify as a formula.

The paper was a safety paper. Its authors, several of whom later founded Anthropic, were looking for a way to make systems do what people meant rather than what they had literally been told, the problem Bostrom had made famous.

Three years later the same recipe was applied to language: a model writes two summaries, a person picks the better one, and the model is trained towards the preference. That became reinforcement learning from human feedback, the step that turned GPT-3 into InstructGPT and InstructGPT into ChatGPT. The chatbot that talks to you politely was tuned by the method in this paper.

what it led to · 32 events downstream, through 2026

Built on it directly:

  1. 2020Learning to summarise from human feedbackV
  2. 2022InstructGPTVI
  3. 2026A model escapes its sandboxVII

And, through them, by era:

sources · 1

See this era in the exhibition →Back to the timeline