Skip to content
The Shape of Intelligence

Boosting: weak learners made strong

Robert Schapire proves that any learner slightly better than chance can be combined into one as accurate as you like; ensembles become a science.

category
theory
significance
2 of 5
people
Robert Schapire, Yoav Freund
organisations
Massachusetts Institute of Technology, AT&T Bell Laboratories

what had to happen · 0 events back to 1943

A root. Nothing in the archive precedes it.

Michael Kearns and Leslie Valiant had asked, in the framework of computational learning theory, whether a rule that is only slightly better than guessing could be turned into one that is very good. Robert Schapire's 1990 paper answered yes, constructively: train a weak learner, train another on the examples the first gets wrong, train a third on the disagreements, and combine them. Repeating the trick drives the error down as far as you like. Yoav Freund and Schapire's AdaBoost of 1995 made the procedure practical, reweighting the data rather than resampling it, and the two shared the Gödel Prize for it in 2003.

Boosting was the most successful learning method of the 1990s that was not a neural network. It made weak, cheap models, usually shallow decision trees, into strong ones, and it came with a theory that explained why it rarely overfit. Viola and Jones's boosted face detector of 2001 put it in every digital camera.

Gradient boosting, its extension by Jerome Friedman in 2001, became XGBoost and LightGBM, which won most structured-data competitions of the 2010s and still run a large part of the world's fraud detection, credit scoring and advertising.

what it led to · 0 events downstream

A leaf, for now. Nothing in the archive has built on it yet.

sources · 1

See this era in the exhibition →Back to the timeline