Skip to content
The Shape of Intelligence

InstructGPT

OpenAI fine-tunes GPT-3 with human feedback to follow instructions; a model a hundred times smaller is preferred by people to the original, and RLHF becomes the standard.

category
model
significance
4 of 5
people
Long Ouyang, Jeff Wu, Jan Leike, Paul Christiano
organisations
OpenAI

what had to happen · 45 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

GPT-3 continued text. Asked to explain the moon landing to a six-year-old, it might produce a list of other things to explain to six-year-olds, because that is a plausible continuation. InstructGPT, announced on 27 January 2022, was the same model taught to do what it was asked. Forty contractors wrote demonstrations of good responses to prompts from the API; the model was fine-tuned on them; the contractors then ranked the model's outputs, a reward model learned the rankings, and the model was optimised against the reward with a leash to keep it from drifting.

Labellers preferred the responses of the 1.3-billion-parameter InstructGPT to those of the 175-billion-parameter GPT-3. The model made up facts less often, was less toxic, and followed instructions it had never seen in training, in languages that were barely present in the feedback. The cost of the alignment step was a tiny fraction of the cost of pre-training.

InstructGPT became the default model of the API in January 2022, and a sibling trained the same way on conversation became ChatGPT ten months later. The recipe has been used on every chat model since. The instrument on this page lets you be the labeller.

what it led to · 30 events downstream, through 2026

Built on it directly:

  1. 2022ChatGPTVI
  2. 2023ClaudeVI
  3. 2023GPT-4VI

And, through them, by era:

sources · 2

See this era in the exhibition →Back to the timeline