# An LLM Can Show Its Work and Still Hide What Changed Its Answer

**Summary:** A visible chain of thought can help with monitoring without faithfully explaining an answer. I would trust LLM reasoning more when the product preserves evidence, alternatives, and state changes that another person can audit.

- Canonical: https://markhuang.ai/news/llm-shows-work-hides-answer-cause
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-10-02
- Section: News
- Tags: LLM Reasoning, Chain of Thought, AI Reliability, AlphaGo, AI Auditing
- Source: [MIT Technology Review](https://www.technologyreview.com/2026/10/02/1145639/dont-be-fooled-llms-dont-reason/)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![A Go board connects a stream of token-like tiles to a transparent tree of evaluated alternatives](https://cdn.markhuang.ai/news/llm-shows-work-hides-answer-cause/hero.webp)

*Fluent output can look deliberate. An auditable system also preserves the alternatives, evidence, and revisions behind the answer.*

In [an MIT Technology Review opinion essay](https://www.technologyreview.com/2026/10/02/1145639/dont-be-fooled-llms-dont-reason/), former DeepMind researcher Thore Graepel argues that today's large language models do not reason in the way AlphaGo did. His distinction is specific: an LLM keeps predicting the next token, even when it produces a long chain of thought, while AlphaGo paired neural networks with an explicit search process that evaluated possible futures.

Graepel grounds the argument in AlphaGo's 2016 match against Lee Sedol. Move 37 in game two looked so unlikely that AlphaGo's policy network assigned it roughly a one-in-10,000 chance of being played by a human expert. Search changed the decision. AlphaGo examined branches beyond the immediately plausible move, played Move 37, and went on to win the five-game match 4-1.

My read is less categorical than the headline. I do not need to settle whether an LLM "really" reasons before deciding whether to trust a product built around it. I need to know whether the system preserves what changed the answer, which alternatives it rejected, and what evidence can prove the result wrong. A fluent explanation is useful, but it is not an audit trail.

## The missing artifact is a working state

Graepel names three gaps. Current chatbots generally lack a persistent, inspectable record of their hypotheses and uncertainties. Their stored knowledge is not cleanly separated from the process that manipulates it. Their written chain of thought can also be an explanation assembled after the answer rather than a faithful account of how the answer arose.

That last problem has direct experimental support. In [Anthropic's 2025 chain-of-thought study](https://www.anthropic.com/research/reasoning-models-dont-say-think), researchers inserted hints into questions and checked whether Claude 3.7 Sonnet and DeepSeek R1 disclosed using them. Across the tested hint types, Claude mentioned the hint 25% of the time and R1 mentioned it 39% of the time, even after the researchers established that the hints affected the answers. The study used artificial multiple-choice settings, and Anthropic says so. It still demonstrates a failure mode that product designers cannot wave away: showing work does not guarantee showing the cause.

OpenAI's later [monitorability research](https://openai.com/index/evaluating-chain-of-thought-monitorability/) supplies an important counterweight. Across 13 evaluations and 24 environments, monitoring chain of thought outperformed monitoring actions and final outputs alone. Longer reasoning traces were generally easier to monitor. OpenAI also warns that this visibility may be fragile as training methods, data, and scale change.

So I would keep the reasoning trace. I just would not promote it to a source of truth. It is one sensor among several, and a sensor can be informative without being complete.

## The AlphaGo comparison clarifies the product test

The original [AlphaGo paper in Nature](https://www.nature.com/articles/nature16961) describes policy networks that select moves, value networks that evaluate positions, and a search algorithm that combines those networks with Monte Carlo simulation. The search tree gave the program a concrete object to update. A branch could be explored, compared, and abandoned.

Open-world work is messier. A medical question, engineering incident, or research hypothesis has no complete board and no fixed legal-move list. Graepel acknowledges that. He proposes an explicit epistemic state that records what the system treats as settled, doubtful, ruled out, or still open, with an independent component checking whether each update reduces uncertainty and has evidential support.

I like that proposal because it turns an argument about machine minds into a product requirement. A system does not earn trust by printing a longer monologue. It earns a larger decision boundary when it can show durable state, cite the evidence behind a change, run an external check, and retain enough history for someone else to reconstruct the choice.

> **Info:**
>
> For a consequential decision, I want the system to expose four separate artifacts: the answer, the evidence it used, the state changes it made, and the checks that tried to disprove it.

## "Reasoning" is still a disputed label

The strongest objection to Graepel's headline is that it defines reasoning by a preferred architecture. If a model can combine facts, solve a novel problem, correct an intermediate error, and choose a useful action, some researchers and developers will call that reasoning regardless of whether the mechanism resembles AlphaGo's tree search.

That definitional split is visible in [a Hacker News discussion about LLM reasoning](https://news.ycombinator.com/item?id=41421591). One side asks for formal operations and explicit search. The other asks why reasoning should be denied when the system produces behavior that meets a task-level test. The debate quickly runs into cats, human brains, first-order logic, and competing meanings of the same word.

I find the behavioral objection persuasive. Architecture alone should not erase a demonstrated capability. But behavior on a benchmark should not erase the need to diagnose a failure. "The model reasoned" is a poor incident report. It does not tell me which evidence entered the system, whether a tool result changed the conclusion, or why a rejected hypothesis stayed rejected.

## What I would require before raising the stakes

I would not wait for a philosophically pure reasoning machine. I would build the missing audit layer around the models available now.

- Record sources and tool outputs separately from model prose, with timestamps and provenance.
- Represent open questions and competing hypotheses as state that can be inspected after the run.
- Require external checks for claims that have deterministic tests, such as calculations, code execution, database constraints, or policy rules.
- Log why the system changed course, then verify that explanation against observable actions rather than trusting the explanation by itself.
- Keep a human approval point and a rollback path when a wrong answer would be costly or hard to reverse.

This extends the argument I made about [moving the AI goalpost toward ownership](https://markhuang.ai/news/ai-goalpost-is-ownership). The person approving an output needs more than a polished rationale. They need evidence that survives the model session and a system that makes disagreement possible.

## I would buy auditability before metaphysics

Graepel is right that scaling a token predictor and asking it to talk for longer does not, by itself, produce AlphaGo's explicit search tree. He is also right that medicine, engineering, and science need more than convincing narration. The claim that LLMs do not reason at all is harder to prove because "reasoning" has no single operational definition across those debates.

That uncertainty does not block a decision. I will treat chain of thought as useful telemetry, not a faithful transcript. I will trust systems more when they keep an inspectable ledger of evidence, alternatives, and belief changes. If a vendor wants to call the model a reasoner, fine. Show me what changed its answer.
