Skip to content
The Shape of Intelligence

· turning point

Word2vec

Mikolov's team at Google learns word vectors from billions of words in hours, and shows that king − man + woman ≈ queen; meaning becomes arithmetic.

category
theory
significance
5 of 5
people
Tomáš Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean
organisations
Google

what had to happen · 10 events back to 1943

Every event this one built on, transitively, in order. Direct influences are marked.

Bengio's 2003 model had learned word vectors as a by-product and had been too slow to train on much. Tomáš Mikolov's insight, published from Google on 16 January 2013, was to throw away the neural network and keep the vectors. Word2vec trains a word's vector to predict the words around it, or the other way round, with a model so simple, a single linear layer, that it processes billions of words in a day on a desktop. The vectors it learns are as good as the slow ones or better.

The paper's famous result is that the vectors carry structure nobody put in. Subtract the vector for "man" from "king", add "woman", and the nearest word to the result is "queen". Paris minus France plus Italy gives Rome. Relations that linguists describe were, in the space the model had learned from co-occurrence alone, straight lines.

Word2vec made embeddings the standard representation of text within a year, and the idea generalised: sentences, images, proteins, products and users all became vectors trained to predict their context. The transformers of 2017 begin with an embedding layer that is word2vec's direct descendant, and the vector databases of the 2020s search by the same arithmetic. The instrument on this page runs the real vectors.

what it led to · 0 events downstream

A leaf, for now. Nothing in the archive has built on it yet.

sources · 2

See this era in the exhibition →Back to the timeline