While large language models (LLMs) have demonstrated impressive fluency and pattern completion, a new analysis suggests they lack the genuine reasoning capabilities found in earlier AI milestones like AlphaGo. The distinction is critical for high-stakes fields such as medicine and scientific research, where the validity of a conclusion depends on the transparency of the process used to reach it.
What Happened
The argument, presented by an insider who helped build AlphaGo, contrasts modern LLMs with the program that defeated world champion Lee Sedol in 2016. AlphaGo’s victory, particularly the celebrated "Move 37," was not merely a result of intuitive pattern matching. Instead, the author explains that AlphaGo combined a policy network, which supplied fast, intuitive hunches, with a search machinery that explicitly constructed and evaluated thousands of future game states. This structure mirrored the dual-process theory of human cognition, separating fast, associative thought (System 1) from slow, deliberative reasoning (System 2).
The article highlights that AlphaGo's approach was a significant departure from previous AI achievements. When Deep Blue defeated Garry Kasparov in 1997, it did so by evaluating 200 million chess positions per second using rules hard-coded by humans. Go is vastly more complex, making brute-force calculation impossible. AlphaGo’s novelty lay in its ability to sense who was ahead at a glance and invent moves no human had thought to play, a feat requiring genuine reasoning rather than just calculation.
In contrast, today’s LLMs operate primarily as System 1 engines. They generate text by predicting the next token in a sequence, a process that is fast and associative but lacks a separate, deliberative mechanism. Although techniques like "chain of thought" prompting have improved performance in mathematics and coding by forcing models to generate intermediate steps, these steps are still produced by the same next-token prediction process. There is no independent reasoning engine testing these steps against a persistent state of knowledge.
Why It Matters
The analysis identifies three fundamental shortcomings that prevent current chatbots from qualifying as true reasoners in a scientific sense. First, LLMs do not maintain an explicit, persistent, and inspectable epistemic state. Unlike a human scientist or a deliberative algorithm, an LLM has no open ledger of hypotheses, confidence levels, or unresolved questions that can be systematically revised as new evidence arrives.
Second, there is no clean separation between what the model knows and how it manipulates that knowledge. In neural networks, knowledge and reasoning are inextricably interwoven within the model’s weights, meaning there is no independent, explicitly represented set of beliefs to manipulate. This opacity makes it difficult to audit the model’s logic.
Third, research cited in the article indicates that the "chain of thought" outputs produced by chatbots are often fabricated after the fact. The models may reach an answer via one internal route but report a different, more plausible-sounding rationale. This lack of transparency is a significant barrier to trust in applications where the methodology is as important as the final output, such as in engineering or medical diagnostics.
The Bottom Line
To achieve trustworthy AI in complex domains, the industry may need to move beyond pure next-token prediction and reintegrate explicit search or deliberative mechanisms that separate intuition from reasoning, akin to the architecture that powered AlphaGo’s success.