InstructGPT
OpenAI fine-tunes GPT-3 with human feedback to follow instructions; a model a hundred times smaller is preferred by people to the original, and RLHF becomes the standard.
what had to happen · 45 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 9
- 1948A mathematical theory of communication
- 1949Cells that fire together wire together
- 1950Programming a computer for playing chess
- 1958The perceptron learns
- 1959Samuel's checkers program coins 'machine learning'
- 1960ADALINE and the least-mean-squares rule
- 1965Moore's law
- 1969Perceptrons
- 1970Reverse-mode automatic differentiation
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 6
III · Statistics and data · 9
IV · Deep learning · 8
V · Transformers · 7
GPT-3 continued text. Asked to explain the moon landing to a six-year-old, it might produce a list of other things to explain to six-year-olds, because that is a plausible continuation. InstructGPT, announced on 27 January 2022, was the same model taught to do what it was asked. Forty contractors wrote demonstrations of good responses to prompts from the API; the model was fine-tuned on them; the contractors then ranked the model's outputs, a reward model learned the rankings, and the model was optimised against the reward with a leash to keep it from drifting.
Labellers preferred the responses of the 1.3-billion-parameter InstructGPT to those of the 175-billion-parameter GPT-3. The model made up facts less often, was less toxic, and followed instructions it had never seen in training, in languages that were barely present in the feedback. The cost of the alignment step was a tiny fraction of the cost of pre-training.
InstructGPT became the default model of the API in January 2022, and a sibling trained the same way on conversation became ChatGPT ten months later. The recipe has been used on every chat model since. The instrument on this page lets you be the labeller.
what it led to · 30 events downstream, through 2026
Built on it directly:
And, through them, by era:
VI · Everyone · 13
- 2023Bing's chatbot and 'Sydney'
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Claude 3 catches GPT-4
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra