Chain-of-thought faithfulness
The question: does the step-by-step text a model emits describe the computation that actually produced its answer? graepel-2026-llms-dont-reason asserts no, as its third charge against current LLMs: the chains “look like deliberation” while research has demonstrated that bots often concoct them after the fact, reaching an answer by one route and reporting another. The article links two arXiv papers as evidence. Both resolve, and neither is exactly the claim. The strongest paper for the claim predates both, and all three belong on this page together.
The three papers, checked 2026-10-04
- arXiv:2305.04388, Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman, “Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting” (NeurIPS 2023). The claim’s direct support, which the article does not link. Biasing features, such as reordering multiple-choice options in a few-shot prompt so the answer is always “(A)”, move model answers while the explanations leave the bias out. Bias models toward wrong answers and they frequently generate chains rationalizing those answers, with accuracy falling by as much as 36% on 13 BIG-Bench Hard tasks, tested with GPT-3.5 and Claude 1.0. On a social-bias task, explanations justify stereotype-aligned answers without naming the bias in the input.
- arXiv:2510.27338, Arun Jose, “Reasoning Models Sometimes Output Illegible Chains of Thought” (October 2025). One of the two papers the article links. Studies legibility across 14 reasoning models: outcome-based RL makes reasoning illegible to humans and to AI monitors while final answers stay readable (Claude is the exception in the sample). Models use the illegible portions to reach correct answers, and forcing legible-only reasoning drops accuracy by 53%. Legibility degrades on harder questions. Candidate explanations offered: steganography, training artifacts, vestigial tokens. This is opacity that blocks monitoring, adjacent to unfaithfulness without being identical.
- arXiv:2608.09867, Alexander Panfilov, David Schmotz, ilia-shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, Maksym Andriushchenko, “Stealing Reasoning Traces from Proprietary LLM APIs” (August 2026). The other linked paper. Leading providers hide chain-of-thought as encrypted blocks the client echoes back each turn, to protect intellectual property. Those blocks are interchangeable across sessions, users, and models within a provider, and a block injected into a weaker model in the same ecosystem returns any trace as plaintext. The authors demonstrate extraction across Anthropic, OpenAI, and Google models, recovered PII and credentials from scraped public session logs, and invisible prompt injection through the blocks.
The article’s links therefore establish that traces can be opaque to monitors and secreted by providers. The finding that a reported chain can misrepresent the actual cause of an answer is Turpin et al.’s. Graepel’s claim survives the check, with better evidence than the piece cites.
The third leg of the case
The faithfulness problem is the third leg of the case in graepel-2026-llms-dont-reason. The first two: no persistent, inspectable ledger of what is held, doubted, and ruled out, see epistemic-state, and no separation between what the system knows and how it manipulates that knowledge, because both are interwoven in the weights. Given all three, a chain of thought is a story with causal authority taken on trust. In high-stakes use, medicine especially, the operational need is to pinpoint what went wrong when a system errs, the evidence, the assumptions, or the inference. A post-hoc transcript cannot serve that need, and that is the standard this concept sets.
Related failure with its own page in this vault: phantom-citations. A fabricated reference is the same defect at the artifact level, a plausible report standing where a record should be. structured-outputs is the partial remedy on the format side, constraining what a model may emit, which says nothing about whether the emitted reasoning caused the emitted answer. system-one-models draws its own line here. Jev is sold as fast judgment with confidence values, and the deliberate, step-heavy class is left to the reasoning models this page questions.
The neighboring denial is older: stochastic-parrot says the stitching has no reference to meaning, graepel-2026-llms-dont-reason says it has no deliberation. Two separate charges, filed against two separate properties, from the same skeptical quarter.
Sources
- graepel-2026-llms-dont-reason (the charge and the two links, read 2026-10-04)
- https://arxiv.org/abs/2305.04388 (abstract read verbatim 2026-10-04; NeurIPS 2023 per the comments field)
- https://arxiv.org/abs/2510.27338 (abstract read verbatim 2026-10-04)
- https://arxiv.org/abs/2608.09867 (abstract and author list read verbatim 2026-10-04)