o1 and reasoning models
OpenAI trains a model to think before it answers, spending more compute at inference on a hidden chain of thought; a second scaling axis opens and mathematics falls.
what had to happen · 53 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 11
- 1948A mathematical theory of communication
- 1949Cells that fire together wire together
- 1950Programming a computer for playing chess
- 1950Computing machinery and intelligence
- 1958The perceptron learns
- 1959Samuel's checkers program coins 'machine learning'
- 1960ADALINE and the least-mean-squares rule
- 1965Moore's law
- 1966ELIZA
- 1969Perceptrons
- 1970Reverse-mode automatic differentiation
W1 · The first winter · 2
II · Connection · 3
W2 · The second winter · 6
III · Statistics and data · 9
IV · Deep learning · 9
V · Transformers · 8
- 2017Attention is all you need
- 2017Deep reinforcement learning from human preferences
- 2017AlphaGo Zero learns from nothingdirect
- 2018GPT: generative pre-training
- 2019GPT-2 and the model too dangerous to release
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Learning to summarise from human feedback
VI · Everyone · 4
- 2022InstructGPT
- 2022Chain-of-thought promptingdirect
- 2022ChatGPT
- 2023GPT-4direct
Chain-of-thought prompting had shown that models did better when they wrote out their reasoning. o1, previewed on 12 September 2024, was trained to do it. Using reinforcement learning on problems with checkable answers, mathematics, code, science, the model learned to produce long private chains of thought, to try approaches, notice errors and back up, before giving a reply. On a qualifying exam for the mathematical olympiad it solved 83 percent of problems where GPT-4o solved 13; on competitive programming it reached the 89th percentile of human entrants.
The company's chart showed accuracy rising smoothly with the compute spent at inference, on the same logarithmic axes as the 2020 scaling laws. Training compute had been the lever for a decade; now there was a second one, and a model could be made smarter by letting it think longer. The chain of thought was hidden from users, a choice OpenAI defended on safety and competitive grounds.
DeepSeek showed in January 2025 that the method could be reproduced cheaply and openly, and every laboratory shipped reasoning models within months. The gold-medal olympiad results of July 2025 came from their descendants.
what it led to · 8 events downstream, through 2026
Built on it directly:
- 2025DeepSeek-R1VII
- 2025Gold at the Mathematical OlympiadVII
- 2025GPT-5VII
And, through them, by era: