Skip to content
The Shape of Intelligence

GPT-2 and the model too dangerous to release

OpenAI's 1.5-billion-parameter model writes coherent pages of text from a prompt; the lab withholds the full weights over misuse fears, and the argument about openness begins.

category
model
significance
4 of 5
people
Alec Radford, Jeffrey Wu, Dario Amodei, Ilya Sutskever
organisations
OpenAI

what had to happen · 31 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

GPT-2 was GPT scaled ten times, to 1.5 billion parameters, and trained on WebText, eight million pages linked from Reddit posts with at least three upvotes. Given a prompt it continued in the same style for paragraphs, inventing a news story about unicorns in the Andes that was quoted everywhere. It also answered questions, translated and summarised without being trained for any of those tasks, simply because the tasks appeared, in some form, in the text it had read. OpenAI called this zero-shot task transfer, and the finding that abilities emerged from scale alone set the agenda for the next five years.

The release, on 14 February 2019, was staged. Citing the risk of automated disinformation, OpenAI published only the smallest of four models and released the rest over nine months as it studied misuse. Critics called it a publicity stunt; supporters called it the first responsible-disclosure process for a model; both were partly right, and the full model, when it came, caused no visible harm.

GPT-2 also fixed the tokeniser. Its byte-pair encoding, which splits text into about 50,000 sub-word pieces, is what the tokens instrument on this site shows, and its successors in GPT-3, GPT-4 and beyond use the same scheme.

what it led to · 48 events downstream, through 2026

Built on it directly:

  1. 2020Scaling laws for neural language modelsV
  2. 2020GPT-3V
  3. 2020Learning to summarise from human feedbackV

And, through them, by era:

sources · 2

See this era in the exhibition →Back to the timeline