AlphaGo Defeats Lee Sedol

3 minute read

Published:

AlphaGo, developed by DeepMind (acquired by Google in January 2014), defeated Lee Sedol — a 9-dan professional Go player widely considered the best active player of the 2010s — in a five-game match held March 9–15, 2016 in Seoul, South Korea, by a score of 4–1. The result shocked the professional Go community: Go’s enormous complexity (roughly 10^170 possible board positions versus chess’s estimated 10^44) had led experts to estimate that computers would not reach top-human level for another decade. AlphaGo had already defeated European Go champion Fan Hui 5–0 in October 2015, but Fan Hui was ranked approximately 700th in the world at the time; Lee Sedol was a fundamentally different caliber of opponent. The match was broadcast live on YouTube and watched by an estimated 200 million viewers in China and internationally. AlphaGo’s architecture combined two deep convolutional neural networks: a policy network (trained first by supervised learning on approximately 30 million positions from the KGS Go server, then refined by reinforcement learning through self-play) that predicted the probability distribution over next moves, and a value network (trained by self-play) that estimated the probability of winning from a given board position. These networks guided a Monte Carlo Tree Search that explored the game tree selectively, simulating future positions along the most promising branches rather than searching exhaustively.

Move 37 in game two became the most discussed moment in the match’s first three games. AlphaGo played a shoulder hit on the fifth line on the right side of the board — a placement that Go commentators broadcasting the game described as either a mistake or inexplicable, since top human players almost never play in that location at that stage of a game. Korean commentator Michael Redmond, a 9-dan professional, said it was “a very strange move” initially; it took several minutes of on-air analysis before the professional viewers recognized that the move was not a mistake but an extremely far-sighted strategic maneuver. AlphaGo’s own internal data later indicated it estimated the probability of a human playing Move 37 at 1 in 10,000 — not because it was a mistake, but because it was a type of move that trained human intuition deprioritizes. Lee Sedol said afterward that he had “never thought of a human playing Move 37,” and that it had shaken his understanding of what Go positions were worth pursuing. AlphaGo won games 1, 2, and 3 consecutively, taking the match regardless of the remaining results.

Lee Sedol’s single victory in game four became equally celebrated for technical reasons. In a complex position, Lee played Move 78 — a wedge into the center of the board that created a ko fight (a type of repetitive capture sequence Go rules limit through the ko rule). AlphaGo’s evaluation network, which had been trained primarily on positions where ko fights were resolved conventionally, responded with moves that were later analyzed as suboptimal under the unusual ko conditions Lee had created — effectively finding a blind spot in the learned evaluation function. Afterward, Lee said the game had been the “happiest moment of my life” despite the 4–1 match loss. DeepMind’s research team acknowledged that the ko situation exposed a genuine evaluation weakness. The match accelerated investment in deep reinforcement learning beyond games: it demonstrated that neural networks trained through self-play could exceed human expert performance on tasks where human expertise had been considered unreachable. AlphaGo Master (December 2016–January 2017) went 60–0 against top online professionals under tournament conditions; AlphaGo Zero (October 2017) was trained entirely from self-play without any human game data and was stronger than all previous AlphaGo versions; and AlphaZero (December 2017) generalized the same algorithm to chess and shogi, achieving superhuman performance in both from scratch in under 24 hours of training.