AlphaGo

DeepMind’s Go program, built by a team including thore-graepel. In March 2016 it beat Lee Sedol, then reigning world-class professional and one of the greatest players of all time, by four games to one in a five-game match in Seoul. Its 2015 match win against Fan Hui is reported on Wikipedia. The architecture is central to graepel-2026-llms-dont-reason, where thore-graepel uses it as the existence proof that machine reasoning, in his sense, was already demonstrated in 2016 and is absent from today’s LLMs.

Architecture, as the article describes it

Two systems. A policy network trained to guess what move a strong human would play, the intuitive part. Search machinery on top: it explicitly constructs and searches a game tree with thousands of branches, each branch a possible future, and weighs the consequences of proposed moves past immediate plausibility. The tree also stores the system’s record of what it knows about the position, each move and position annotated with the neural networks’ judgments. That record is the model for epistemic-state.

The primary paper is Silver et al., “Mastering the game of Go with deep neural networks and tree search,” Nature 529:484-489 (2016). The successors, AlphaGo Zero, AlphaZero, and MuZero, are named on thore-graepel’s site.

Move 37

Game two, move 37: a stone on the fifth line that looked like a gift to the human opponent. Commentators suspected a glitch. The policy network regarded the move as nothing special, roughly a one-in-10,000 chance of an expert human playing it. The search machinery chose it anyway. Lee Sedol’s reaction, quoted in the article: “I thought AlphaGo was based on probability calculation and that it was merely a machine. But when I saw this move, I changed my mind. Surely, AlphaGo is creative.”

Graepel’s reading of the episode is that the public one is wrong: move 37 was not machine intuition. It was deliberation overriding intuition, the half of the architecture that today’s next-token prediction does not have.

Sources