Epistemic state

A named requirement from graepel-2026-llms-dont-reason. A reasoning system, in Thore Graepel’s sense, keeps an explicit, persistent, inspectable record of where it stands: what it holds as settled, what it doubts, what it has ruled out, and which questions stay open. The record gets revised systematically as new information arrives. Chain-of-thought text is not this. It is a transcript, produced once, of a generation that already happened.

The template is alpha-go’s game tree. The tree stores every variation the program considered, each move and position annotated with judgments from its neural networks. Reasoning proceeds by updating the tree, and the final move is synthesized from what the tree contains. Graepel generalizes the structure: reasoning is a sequence of moves that change the epistemic state to advance knowledge and reduce uncertainty. The moves are deducing consequences, breaking problems into parts, and deciding what question to ask, what calculation to perform, or what experiment to run next. That last one, choosing the next probe, is where open-world reasoning exceeds board games: the state of affairs is only partially known, the action set is large and variable, and consequences are stochastic or unknown.

Two more parts complete the design. An independent evaluator judges each move by how much it actually resolves uncertainty, and beliefs update only when the change is backed by evidence. With those rules enforced, the system accumulates certified knowledge and improves its reasoning policy by learning from past reasoning episodes. Graepel’s phrase for the assembled thing is “the scientific method on steroids, with the purpose of producing knowledge that can withstand scrutiny.”

The three denials

The article denies three things, detailed in chain-of-thought-faithfulness: an LLM does not maintain such a ledger, does not separate stored knowledge from its manipulation because both live interwoven in the weights, and its chain of thought is often a post-hoc report. The positive claim is that all three failures share one cure, and the cure is the ledger.

Adjacents in this vault

system-one-models is the opposite bet, drawn in the open: Jev returns typed output with a confidence value per judgment, which is inspectable, but state does not carry between calls, so nothing accumulates between them. TypeSafe’s own docs assign long-running deliberation to reasoning models, the class Graepel says cannot deliberate. structured-outputs matters here because a ledger is only useful if its entries have shapes that can be checked. The connection between the two sources is this vault’s observation. Neither names the other.

Sources