Skip to content
The Shape of Intelligence

Finding structure in time

Jeffrey Elman's recurrent network feeds its own hidden state back as input and learns grammar-like structure from sequences of words with no labels.

category
theory
significance
3 of 5
people
Jeffrey Elman
organisations
University of California San Diego

what had to happen · 8 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

Jeffrey Elman was a linguist, and the question in his 1990 paper was whether a network could learn anything about language from exposure alone. His network was simple: a standard hidden layer whose activations were copied, at each step, into a set of context units that fed back in at the next step, so that the network's state carried a memory of what it had seen. He trained it to predict the next word in sentences generated from a small grammar.

It could not predict the exact word, because that is not predictable. What it learned instead was the structure. The hidden states clustered nouns apart from verbs, animate from inanimate, and the network's predictions respected agreement across intervening words, without anyone telling it what a noun was. The model had discovered categories from the statistics of sequence.

The paper is the ancestor of every language model that followed, and its title is the programme. Next-word prediction as a task, learned representations as the product, and grammar emerging rather than being written: GPT is Elman's network with a hundred billion times the parameters and attention instead of recurrence. The weakness he found, that memory decayed over long sequences, was named the vanishing gradient the next year and solved by the LSTM in 1997.

what it led to · 78 events downstream, through 2026

Built on it directly:

  1. 1991The vanishing gradient problemW2
  2. 1997Long short-term memoryIII
  3. 2003A neural probabilistic language modelIII

And, through them, by era:

III · Statistics and data · 1
  1. 2010Rectified linear units
IV · Deep learning · 14
  1. 2012AlexNet wins ImageNet
  2. 2013Word2vec
  3. 2013Deep Q-networks play Atari
  4. 2014Google buys DeepMind
  5. 2014Generative adversarial networks
  6. 2014Attention
  7. 2014Sequence to sequence learning
  8. 2015Batch normalisation
  9. 2015Residual networks
  10. 2015OpenAI is founded
  11. 2016AlphaGo beats Lee Sedol
  12. 2016Google reveals the TPU
  13. 2016WaveNet
  14. 2016Google Translate goes neural
V · Transformers · 17
  1. 2017Attention is all you need
  2. 2017Deep reinforcement learning from human preferences
  3. 2017AlphaGo Zero learns from nothing
  4. 2018GPT: generative pre-training
  5. 2018BERT
  6. 2018AlphaFold enters the protein-folding contest
  7. 2019GPT-2 and the model too dangerous to release
  8. 2019The bitter lesson
  9. 2020Scaling laws for neural language models
  10. 2020GPT-3
  11. 2020Learning to summarise from human feedback
  12. 2020An image is worth 16×16 words
  13. 2020AlphaFold 2 solves protein structure prediction
  14. 2021CLIP and DALL·E
  15. 2021On the dangers of stochastic parrots
  16. 2021Anthropic is founded
  17. 2021GitHub Copilot writes code
VI · Everyone · 29
  1. 2022InstructGPT
  2. 2022Chain-of-thought prompting
  3. 2022Chinchilla: the models were undertrained
  4. 2022PaLM
  5. 2022DALL·E 2
  6. 2022Midjourney opens its beta
  7. 2022Stable Diffusion is released
  8. 2022Galactica lasts three days
  9. 2022ChatGPT
  10. 2023Bing's chatbot and 'Sydney'
  11. 2023LLaMA leaks and open weights take off
  12. 2023Claude
  13. 2023GPT-4
  14. 2023'Pause Giant AI Experiments'
  15. 2023Hinton leaves Google to warn about AI
  16. 2023The US executive order on AI
  17. 2023The Bletchley Declaration
  18. 2023OpenAI fires and rehires its chief executive
  19. 2023Gemini
  20. 2024Sora
  21. 2024Claude 3 catches GPT-4
  22. 2024AlphaFold 3
  23. 2024GPT-4o talks
  24. 2024The EU AI Act enters into force
  25. 2024o1 and reasoning models
  26. 2024The Nobel Prizes go to neural networks
  27. 2024Claude learns to use a computer
  28. 2024The Model Context Protocol
  29. 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
  1. 2025DeepSeek-R1
  2. 2025Claude 4 and Claude Code
  3. 2025Nvidia is worth four trillion dollars
  4. 2025Gold at the Mathematical Olympiad
  5. 2025America's AI Action Plan
  6. 2025GPT-5
  7. 2025Gemini 3
  8. 2025MCP is donated to the Agentic AI Foundation
  9. 2026Claude Fable 5 and the Mythos class
  10. 2026GPT-5.6: Sol, Terra and Luna
  11. 2026A model escapes its sandbox
  12. 2026The EU delays its high-risk AI rules
  13. 2026Claude Fable 5.1
  14. 2026GPT-6 Astra

sources · 1

See this era in the exhibition →Back to the timeline