Don’t be fooled—LLMs don’t reason

2026-10-03 · MIT Technology Review

Don't be fooled—LLMs don't reason

AlphaGo's "Move 37": Intuition Plus Reasoning

In March 2016, AlphaGo played the famous "Move 37" against Lee Sedol. The move looked so absurd that commentators thought it was a glitch, but AlphaGo won the game 4-1. Lee Sedol later admitted that AlphaGo was creative.

Unlike Deep Blue, which defeated Garry Kasparov in 1997 using human hard-coded rules and brute-force search, Go is too complex for exhaustive calculation. AlphaGo won because it could assess the board at a glance and invent moves no human had considered.

Many viewed Move 37 as pure machine intuition, but this is a misunderstanding. The creative choice was driven by AlphaGo's powers of reasoning—capabilities today's AI lacks.

The Dual-System Mechanism: AlphaGo's Architecture

AlphaGo consists of two systems:

  • The Policy Network: Trained to guess what a strong human would play. This "intuitive" part viewed Move 37 as unremarkable, giving it a one in 10,000 chance of being played by a human expert.
  • The Search Machinery: This is what actually chose the move. It looked beyond immediate plausibility, weighed future consequences, and explicitly constructed and searched a game tree with thousands of branches representing different futures.

This mirrors the dual-system theory popularized by Daniel Kahneman:

  • System 1: Fast, gut-level, and effortless. AlphaGo's networks supplied the hunches.
  • System 2: Slow, step-by-step, and deliberative. AlphaGo's search tested these hunches against subsequent moves and countermoves.

Neither half works alone. Intuition alone wouldn't have chosen Move 37, and brute-force search would have struggled to sieve through all possibilities.

The Limitations of Today's Large Language Models

Today's AI models work differently. A large language model simply picks the next token repeatedly, acting as System 1—fast, associative pattern completion.

After ChatGPT's debut, it became clear that fluency wasn't enough. The solution was "chain of thought," where models generate intermediate steps before answering. While beneficial in math and coding, this doesn't introduce a genuinely separate reasoning mechanism. The intermediate reasoning is still produced by the same next-token prediction process, just iterated longer.

Three Shortcomings Preventing LLMs from Reasoning

From a scientific perspective, chatbots fail to qualify as reasoning due to three main flaws:

1. No Persistent Epistemic State: Models lack an open ledger to track hypotheses, confidence levels, evidence, and unresolved questions that should be systematically revised with new information.

2. No Clean Separation of Knowledge and Manipulation: Knowledge and reasoning are inextricably interwoven in the neural network's weights, with no independent set of explicitly represented beliefs.

3. Post-Hoc Fabrication: Research shows that bots often concoct chains of thought after the fact, reaching an answer by one route but reporting another.

The Need for Genuine Reasoning in High-Stakes Fields

In high-stakes applications like medicine, engineering, and scientific research, how a system arrives at a conclusion is as important as the conclusion itself. When a medical diagnosis fails, we must pinpoint whether the reasoning was flawed, the evidence invalid, or the assumptions incorrect.

This is why the author left Google DeepMind. We need a fresh approach to machine reasoning inspired by AlphaGo. AlphaGo maintains a record of what it knows: the game tree. This data structure contains all considered variations and possible futures, with each move annotated by its neural networks. As reasoning progresses, AlphaGo updates the game tree and synthesizes the information to make a decision.

Similarly, for general reasoning, a system should maintain an epistemic state representing what is settled, doubted, ruled out, and still open. Reasoning can then be understood as a sequence of operations updating and querying this state, paving the way for trustworthy and genuinely novel AI insights.

Source