Skip to content
The Shape of Intelligence

· turning point

Attention

Bahdanau, Cho and Bengio let a translation model look back at every source word and learn which to weigh; the mechanism at the heart of the transformer appears.

category
theory
significance
5 of 5
people
Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio
organisations
Université de Montréal, Jacobs University Bremen

what had to happen · 12 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

The encoder–decoder models of 2014 compressed a whole source sentence into one vector and then generated the translation from it. That worked for short sentences and degraded for long ones, because a fixed vector cannot hold forty words. Dzmitry Bahdanau's fix, posted on 1 September 2014 with Kyunghyun Cho and Yoshua Bengio, was to let the decoder look back. At each output word the model computes a weight for every source word, takes a weighted average of their encodings, and uses that as its context. The weights are learned, and when you plot them, they draw the alignment between the two languages.

The paper called it soft alignment; the community called it attention, and the name stuck. It was the first mechanism that let a network choose what to look at, and it removed the bottleneck that had limited sequence models since Elman.

Three years later, Google researchers asked what happened if attention was the only mechanism, with no recurrence at all, and the transformer was the answer. Every large language model computes, at every layer, a version of the weighted average Bahdanau introduced here. The instrument on this site shows the weights a real model computes over a sentence you type.

what it led to · 57 events downstream, through 2026

Built on it directly:

  1. 2016Google Translate goes neuralIV
  2. 2017Attention is all you needV

And, through them, by era:

V · Transformers · 12
  1. 2018GPT: generative pre-training
  2. 2018BERT
  3. 2019GPT-2 and the model too dangerous to release
  4. 2020Scaling laws for neural language models
  5. 2020GPT-3
  6. 2020Learning to summarise from human feedback
  7. 2020An image is worth 16×16 words
  8. 2020AlphaFold 2 solves protein structure prediction
  9. 2021CLIP and DALL·E
  10. 2021On the dangers of stochastic parrots
  11. 2021Anthropic is founded
  12. 2021GitHub Copilot writes code
VI · Everyone · 29
  1. 2022InstructGPT
  2. 2022Chain-of-thought prompting
  3. 2022Chinchilla: the models were undertrained
  4. 2022PaLM
  5. 2022DALL·E 2
  6. 2022Midjourney opens its beta
  7. 2022Stable Diffusion is released
  8. 2022Galactica lasts three days
  9. 2022ChatGPT
  10. 2023Bing's chatbot and 'Sydney'
  11. 2023LLaMA leaks and open weights take off
  12. 2023Claude
  13. 2023GPT-4
  14. 2023'Pause Giant AI Experiments'
  15. 2023Hinton leaves Google to warn about AI
  16. 2023The US executive order on AI
  17. 2023The Bletchley Declaration
  18. 2023OpenAI fires and rehires its chief executive
  19. 2023Gemini
  20. 2024Sora
  21. 2024Claude 3 catches GPT-4
  22. 2024AlphaFold 3
  23. 2024GPT-4o talks
  24. 2024The EU AI Act enters into force
  25. 2024o1 and reasoning models
  26. 2024The Nobel Prizes go to neural networks
  27. 2024Claude learns to use a computer
  28. 2024The Model Context Protocol
  29. 2024DeepSeek-V3 trained for $5.6 million
VII · Agents · 14
  1. 2025DeepSeek-R1
  2. 2025Claude 4 and Claude Code
  3. 2025Nvidia is worth four trillion dollars
  4. 2025Gold at the Mathematical Olympiad
  5. 2025America's AI Action Plan
  6. 2025GPT-5
  7. 2025Gemini 3
  8. 2025MCP is donated to the Agentic AI Foundation
  9. 2026Claude Fable 5 and the Mythos class
  10. 2026GPT-5.6: Sol, Terra and Luna
  11. 2026A model escapes its sandbox
  12. 2026The EU delays its high-risk AI rules
  13. 2026Claude Fable 5.1
  14. 2026GPT-6 Astra

sources · 2

See this era in the exhibition →Back to the timeline