Skip to content
The AI-First Web

AlphaGo veteran says chatbot reasoning is often a story told after the answer

Thore Graepel argues that today's AI shows its steps but keeps no auditable record of what it knew or doubted.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
AlphaGo veteran says chatbot reasoning is often a story told after the answer

AI-generated image for WebPulse. About our images

In brief
  • AlphaGo veteran Thore Graepel argues chatbots lack an inspectable record of what they know and doubt, and their displayed steps are often written after the answer.
  • This matters where AI informs medicine, engineering or research, because a failure needs a traceable cause, not a convincing explanation.
  • Ask vendors whether shown reasoning drives the answer, where evidence is recorded, and who can investigate a wrong result.

An explanation is not an audit trail

Many AI tools now show their working. They list steps, then give an answer. It looks like a person thinking aloud. Thore Graepel, a core member of the AlphaGo team, argues in MIT Technology Review that the resemblance can mislead.

His claim is narrow and worth taking seriously. A visible chain of steps is not the same as a record you can inspect when something goes wrong. For anyone who signs off on AI in medicine, engineering or research, that gap matters.

What Graepel argues

Graepel is chair of machine learning at University College London. He writes that he recently left his position at Google DeepMind to pursue a different approach to machine reasoning. This is one researcher's view. It is not a finding from a study.

His central point is that today's chatbots do something closer to fast intuition than deliberate reasoning. He names three shortcomings. They keep no persistent, inspectable record of what they believe. They do not separate what they know from how they use it. And their displayed steps may not reflect how they reached the answer.

About 1 in 10,000
Chance an expert human would play AlphaGo's move 37
Source: Thore Graepel, MIT Technology Review (October 2, 2026)

How AlphaGo worked, and why it is the contrast

AlphaGo had two parts. The first, a policy network, was trained to guess what a strong human would play. It rated move 37 as unremarkable.

The second part was a search. It built a game tree: a map of thousands of possible futures, each branch scored by the neural networks. The search weighed consequences and chose the move. Graepel credits that search, not intuition, for the famous move.

He links this to Daniel Kahneman's two modes of thought: fast and gut-level, and slow and deliberate. AlphaGo had both. The game tree was also a written record of what the system had considered.

4-1
Final score of AlphaGo against Lee Sedol
Source: Thore Graepel, MIT Technology Review (October 2, 2026)

How a chatbot differs

A large language model predicts the next word fragment, called a token, over and over. Graepel describes that as the fast, intuitive mode. It is good at completing patterns across almost any subject.

Chain of thought is the newer technique. The model writes intermediate steps before it answers. Graepel says the gains are real, above all in mathematics and coding.

But he stresses it is not a separate reasoning mechanism. The steps come from the same next-token process, run for longer. He adds that research has shown chatbots often reach an answer by one route but report another.

Compare that with the game tree. A tree can be read, questioned and corrected. A set of steps written after the fact cannot be checked against what actually happened inside the model.

Why this reaches the boardroom

Graepel makes the stakes concrete. When a diagnosis or treatment goes wrong, you need to pinpoint the fault. Was the reasoning flawed, the evidence invalid, or an assumption wrong? A story told after the fact cannot answer that.

The lesson here is about evidence. A confident explanation is easy to produce. An explanation you can audit is much harder. Organisations that treat the first as the second take on a risk they cannot see.

Graepel's proposed fix is a system that keeps an explicit record of what it treats as settled, what it doubts and what it has ruled out. An independent part would check each step and update beliefs only when evidence supports the change. That is a research direction, not a product on sale.

Questions to put to your team and vendors

First, ask whether the steps an AI tool displays are known to drive its answer, or only to accompany it. Ask for evidence, not reassurance.

Second, ask where the system records its evidence, assumptions and open questions. If nowhere, ask how a wrong conclusion would be traced.

Third, name the person who can investigate a bad AI-assisted decision. Make sure they have something to investigate.

Fourth, match the tool to the task. Graepel says chain-of-thought gains are real in maths and coding. That does not extend automatically to every high-stakes use.

Graepel closes with a standard worth borrowing: conclusions that arise from an auditable sequence of evidence, inference and belief revision, not from a convincing story told afterwards.

200 million
Positions per second Deep Blue evaluated against Kasparov in 1997
Source: Thore Graepel, MIT Technology Review (October 2, 2026)

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: MIT Technology Review.

Share this insight