AlphaGo Prepares for Its Match Against Lee Sedol
Published:
On January 27, 2016, DeepMind published the Nature paper “Mastering the Game of Go with Deep Neural Networks and Tree Search” describing the AlphaGo system and simultaneously announcing that AlphaGo would face Lee Sedol — 9-dan professional Go player, 18-time world champion, and widely considered the strongest player of the 2010s — in a five-game match in Seoul in March 2016. The paper disclosed that AlphaGo had already defeated Fan Hui, the European Go champion, 5–0 in a formal match held secretly at DeepMind’s London offices on October 5–9, 2015 — the first time a computer program had beaten a professional Go player in a formal match on a standard 19×19 board without handicap. Previous Go AIs, including Zen and Crazy Stone, used Monte Carlo Tree Search (MCTS) with handcrafted evaluation functions and required 4-stone or larger handicaps to compete with professional players. Fan Hui, a 2-dan professional and European champion, was ranked approximately 700th in the world — a professional level, but far below top-10 world ranking where Lee Sedol operated. Go’s complexity — a branching factor of approximately 250 legal moves per turn (versus chess’s 35), a game tree estimated at 10^170 positions, and position evaluation that defied the heuristic functions effective in chess — had led the AI community to expect computational professional-level Go was decades away.
AlphaGo’s architecture addressed Go’s complexity through three integrated components. The policy network — a 13-layer convolutional neural network (CNN) with 192 filters per layer — was trained first by supervised learning on approximately 30 million board positions from games played by human experts on the KGS Go server, learning to predict the human expert’s move for each position (achieving 57% prediction accuracy, versus 44% for previous best CNN approaches). This supervised learning initialization gave the policy network a strong baseline for which regions of the game tree were worth exploring. The network was then refined through reinforcement learning: AlphaGo played against earlier versions of itself for thousands of games, updating the policy network’s weights to increase the probability of moves that led to wins, reaching an 80% win rate against the supervised-only version after just 500 iterations of self-play. The value network — also a CNN trained from self-play game positions — predicted the probability of winning from any board position, providing fast position evaluation without requiring complete game simulation to the end. Monte Carlo Tree Search combined both networks: the policy network narrowed MCTS exploration to the most promising moves (reducing the effective branching factor from ~250 to a manageable set), while the value network estimated position values at each leaf node without requiring full rollouts, and traditional MCTS fast rollouts provided additional signal.
The computational resources required for the Lee Sedol-strength version of AlphaGo were substantial: the distributed configuration used by DeepMind in matches deployed across Google’s data center infrastructure used approximately 1,202 CPUs and 176 GPUs. A single-machine version used 48 Google TPU (Tensor Processing Unit) chips, though TPU details were not publicly disclosed at the time. The anticipation created by the Nature paper and the announced Lee Sedol match was significant beyond the Go community: the paper provided the first peer-reviewed documentation that deep reinforcement learning combined with MCTS could reach superhuman level on a game that AI researchers had considered a canonical hard problem. The match was scheduled for March 9–15, 2016, with the prize fund of $1 million going to the winner (donated to charity if AlphaGo won; the International Go Federation and UNICEF were designated recipients). The broader AI community, economists studying automation, and science journalists all focused on the match as a test of whether AlphaGo’s success against Fan Hui would hold at the highest competitive level — and whether the deep learning techniques behind AlphaGo might generalize beyond Go to problems with similar large state spaces and reward signals.
