· turning point
Long short-term memory
Hochreiter and Schmidhuber's memory cell with gates lets recurrent networks learn across a thousand steps; it becomes the engine of speech, translation and text until 2017.
what had to happen · 10 events back to 1943
Every event this one built on, transitively, in order. Direct influences are marked.
00 · One neuron · 1
I · Foundations · 5
W1 · The first winter · 1
II · Connection · 1
- 1986Backpropagation
W2 · The second winter · 2
- 1990Finding structure in timedirect
- 1991The vanishing gradient problemdirect
Hochreiter's 1991 thesis had shown why recurrent networks forgot: the gradient vanished as it travelled back through time. The Long Short-Term Memory, published with Jürgen Schmidhuber in Neural Computation on 15 November 1997, is the fix. Each cell holds a value that passes from step to step unchanged unless a learned gate decides to write to it or erase it, so the error signal can flow back through hundreds of steps without shrinking. The network learns not just what to compute but what to remember.
It took a decade to matter, because it was slow to train and the problems it solved, long sequences, were not yet the problems people had data for. Then, from about 2009, Alex Graves and others showed LSTMs beating everything on handwriting and speech, and by 2015 it was inside Google's speech recogniser, Apple's Siri, Amazon's Alexa and, the following year, Google Translate. Sequence-to-sequence learning and the first attention mechanisms were built on it.
The transformer of 2017 replaced the LSTM in language by removing recurrence altogether, which made training parallel. But for the years in which "neural network" came to mean something that could handle text and sound, the LSTM was the network. It is the most cited neural-network paper of the twentieth century.
what it led to · 59 events downstream, through 2026
Built on it directly:
- 2014AttentionIV
- 2014Sequence to sequence learningIV
- 2016Google Translate goes neuralIV
And, through them, by era:
V · Transformers · 13
- 2017Attention is all you need
- 2018GPT: generative pre-training
- 2018BERT
- 2019GPT-2 and the model too dangerous to release
- 2020Scaling laws for neural language models
- 2020GPT-3
- 2020Learning to summarise from human feedback
- 2020An image is worth 16×16 words
- 2020AlphaFold 2 solves protein structure prediction
- 2021CLIP and DALL·E
- 2021On the dangers of stochastic parrots
- 2021Anthropic is founded
- 2021GitHub Copilot writes code
VI · Everyone · 29
- 2022InstructGPT
- 2022Chain-of-thought prompting
- 2022Chinchilla: the models were undertrained
- 2022PaLM
- 2022DALL·E 2
- 2022Midjourney opens its beta
- 2022Stable Diffusion is released
- 2022Galactica lasts three days
- 2022ChatGPT
- 2023Bing's chatbot and 'Sydney'
- 2023LLaMA leaks and open weights take off
- 2023Claude
- 2023GPT-4
- 2023'Pause Giant AI Experiments'
- 2023Hinton leaves Google to warn about AI
- 2023The US executive order on AI
- 2023The Bletchley Declaration
- 2023OpenAI fires and rehires its chief executive
- 2023Gemini
- 2024Sora
- 2024Claude 3 catches GPT-4
- 2024AlphaFold 3
- 2024GPT-4o talks
- 2024The EU AI Act enters into force
- 2024o1 and reasoning models
- 2024The Nobel Prizes go to neural networks
- 2024Claude learns to use a computer
- 2024The Model Context Protocol
- 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
- 2025DeepSeek-R1
- 2025Claude 4 and Claude Code
- 2025Nvidia is worth four trillion dollars
- 2025Gold at the Mathematical Olympiad
- 2025America's AI Action Plan
- 2025GPT-5
- 2025Gemini 3
- 2025MCP is donated to the Agentic AI Foundation
- 2026Claude Fable 5 and the Mythos class
- 2026GPT-5.6: Sol, Terra and Luna
- 2026A model escapes its sandbox
- 2026The EU delays its high-risk AI rules
- 2026Claude Fable 5.1
- 2026GPT-6 Astra