Graepel 2026, Don’t be fooled, LLMs don’t reason
Opinion piece by thore-graepel, published in mit-technology-review on 2026-10-02, under the
“Don’t be fooled” banner the magazine shares with the Gebru and Bender piece of September. Filed as
text in content/Inbox/2026-10-04-dont-be-fooled-llms-dont-reason.txt at ingest.
The thesis: ten years after alpha-go beat Lee Sedol, today’s AI still lacks the machinery that made that win possible. LLMs are pattern completion, one next token at a time, what the System 1 / System 2 frame of daniel-kahneman calls fast, intuitive thought. Chain of thought does not change the class. It is the same next-token process iterated longer before the model commits. The gains it delivers in math and code are real, and they are sharpened intuition. Deliberation would require a genuinely separate mechanism, and current systems do not have one.
The argument’s parts
The move 37 reading. The public story treats AlphaGo’s move 37, game two against Lee Sedol, as machine intuition. Graepel’s reading, from the team that built it: the policy network regarded the fifth-line move as nothing special, roughly a one-in-10,000 play for an expert human, and the search machinery chose it anyway after weighing thousands of future branches. Intuition alone would not have played it; brute-force search alone could not have sifted Go’s branching. The episode is the existence proof that a machine held a position, weighed futures, and overrode its own instincts.
Three shortcomings. What chatbots do fails as reasoning on three counts, developed on chain-of-thought-faithfulness and epistemic-state: no explicit, persistent, inspectable epistemic state (no ledger of hypotheses, confidences, evidence, open questions); no clean separation between knowledge and its manipulation (both interwoven in the weights); chains of thought concocted after the fact, one route to the answer, another route reported.
The stakes. In medicine, engineering, and science, how a conclusion arrives matters as much as the conclusion. When a diagnosis errs, someone must pinpoint whether the fault was the reasoning, the evidence, or the assumptions. A transcript written after generation cannot answer that.
The prescription. Build the ledger. Maintain an epistemic state holding what is settled, doubted, ruled out, still open. Treat reasoning as moves that change it: deduce, decompose, choose the next question, calculation, or experiment. Let an independent part of the system score each move by realized uncertainty reduction, update beliefs only on evidence, and let the reasoning policy learn from past episodes. His words: “the scientific method on steroids.” LLMs slot in as helpers, suggesting tactics, driving tools, weighing evidence, under that architecture.
The anti-scale line. “I do not think we reach trustworthy machine intelligence by making system 1 bigger. Scale sharpens intuition, but it does not make intuition more deliberative.”
Standing of the claims
This is an opinion piece, first-person, by a participant: he watched move 37 land in Seoul and says he left Google DeepMind over the position argued here. On AlphaGo’s internals his testimony is primary. On the present state of LLMs the piece cites papers rather than data of its own, and the links it carries check out on 2026-10-04 with one nuance: both resolve to work on chains of thought being hidden or illegible (arXiv:2608.09867, arXiv:2510.27338), while the unfaithful-rationalization finding the sentence asserts is Turpin et al. 2023 (arXiv:2305.04388, NeurIPS 2023). All three are filed at chain-of-thought-faithfulness. The open-world-vs-board-game difficulty is conceded in the text itself, so the prescription is a program, not a demo.
Pointers inside the piece
Related stories linked from the article page, un-ingested, MIT Technology Review:
- “When can we say AI made a scientific discovery?” (2026-09-28)
- “AI is rewiring how the world’s best Go players think” (2026-02-27)
- “AI for science needs reasoning, not just data” (2026-08-10)
- “AI’s recursive self-improvement might not come so quickly after all” (2026-08-18), adjacent to zenil-2026-self-improvement-limits
- “Don’t be fooled by this summer of AI hype” (2026-09-22), Timnit Gebru and Emily M. Bender
In this vault
The strongest anti-reasoning position the vault holds, and the first that comes with a concrete architecture rather than a limit theorem. It presses against system-one-models from the other side: TypeSafe’s framing assigns deliberation to reasoning models and sells fast judgment as the complement; Graepel says the complement class cannot deliberate either. Both borrow the Kahneman frame and disagree about who owns System 2. The vault now has the disagreement on file; Ryan’s call what to do with it. The parrot camp sits next door, stochastic-parrot denies the reference side of meaning while this piece denies the deliberation side; different claims, no source in either camp addresses the other’s.
Sources
- https://www.technologyreview.com/2026-10-02/1145639/dont-be-fooled-llms-dont-reason/ (article read in full 2026-10-04)
- https://arxiv.org/abs/2608.09867, https://arxiv.org/abs/2510.27338 (linked papers, abstracts read 2026-10-04)
- https://arxiv.org/abs/2305.04388 (correct referent for the confabulation claim, abstract read 2026-10-04)
- https://vialogue.wordpress.com/2018-02-16/alphago-reflections-and-quotes/ (Lee Sedol quote origin, linked by the article)