I keep getting stuck on the word “reason” in Thore Graepel’s “Don’t be fooled—LLMs don’t reason”.
Graepel helped build AlphaGo. He contrasts its explicit search through possible moves with an LLM’s generation of tokens. He wants reasoning systems to maintain a persistent, inspectable record of beliefs, separate knowledge from the machinery operating on it, and explain how they actually reached their conclusions. I’d like those features in a system I had to rely on. I’m less sure they describe what humans ordinarily do.
Humans do reason—though perhaps not according to Graepel’s definition. That’s why the title deserves a second look. If a definition excludes much of what we ordinarily call human reasoning, we need to decide whether we’re describing a more demanding ideal or have simply misplaced the humans. Similarly, if a definition includes things that don’t fit our instinctive grasp of the concept, that mismatch is worth exploring too.
I ran into something similar in my essay on understanding as compression. Take that definition seriously and you have to ask whether gzip understands a file. My instinct objects. But why, exactly? Following that objection is where the definition starts to get interesting.
Does a regex reason?
Yann LeCun is a deep-learning pioneer and NYU professor who shared the 2018 Turing Award with Geoffrey Hinton and Yoshua Bengio. Endorsing Graepel’s article, he writes, “True reasoning must involve a search,” adding that LLMs lack that capability.
When I read “search and backtracking,” I’m immediately drawn to regexes. Does a regular expression reason? It certainly does the search and backtracking part—or rather, a backtracking regex engine does—but in my gut, that’s not what I mean by reasoning.
That doesn’t refute LeCun. He says search is necessary; he doesn’t say it is sufficient. Graepel’s definition asks for more too. Still, granting the necessary condition leaves quite a bit to explain. What else does the search need to be doing? Would following a logical rule be enough? What about a very simple rule?
I don’t have a satisfying answer yet. Perhaps some regex matching should count as a very limited sort of reasoning. I’d want to understand what that definition buys us before dismissing it just because the example feels wrong.
What if we put the search outside the model? Tree of Thoughts combines an LLM’s proposed steps and evaluations with an external search algorithm. The whole arrangement searches. That doesn’t demonstrate the same search mechanism inside the model, so we need to be clear about which one we’re evaluating. We’ll need the same distinction when a human asks for a notebook.
Would I pass the test?
Apple’s The Illusion of Thinking gives us actual failures to examine. In its puzzle experiments, the tested models benefited from extra thinking at intermediate difficulty, then failed as complexity increased. Giving them the algorithms didn’t eliminate the failures. The revised paper addresses objections about impossible cases and output limits; failures on feasible tasks remain.
I’m not inclined to wave that away. I do want to know what failed. Did the model choose the wrong approach? Lose its place? Misapply a rule? Those are different problems, and “doesn’t reason” doesn’t tell me which one I’m looking at.
Consider the Tower of Hanoi. For a sense of scale, fifteen disks on three pegs require at least 32,767 moves. The recursive algorithm is short enough to fit on a card. Now ask me to write down all the moves, in order, without making a mistake. I’d want a notebook. Better, I’d want to write a program. At some point I’d also want to know why we’re doing this.
Apple didn’t run a matched human comparison, so I can’t tell you how people would fare. The fifteen-disk example is my illustration of the execution problem. Knowing a method and carrying out a long sequence correctly are things we can test separately. It matters whether an error occurred while finding the method or while keeping track of move 8,193.
We do have evidence that the way a puzzle is presented matters to humans. In Experiment 2 of Zhang and Norman’s Tower of Hanoi study, changing which constraints were represented externally changed performance. The same underlying problem can place different demands on the person solving it.
There’s another way a problem can trip us up: change what it’s about. Try four cards.
Four cards show D · F · 3 · 7. Each has a letter on one side and a number on the other.
The rule is: If a card has D on one side, it has 3 on the other.
Which cards must you turn over to check the rule?
D and 7. The D might have the wrong number behind it. The 7 might have a D. You don’t need to turn over the 3: the rule allows other letters to have a 3 too.
Now make the rule: If someone is drinking beer, they must be at least eighteen. You check the beer drinker and the seventeen-year-old. Same logical structure, rather different associations. That doesn’t guarantee everyone will find it easier, but it gives us a way to ask whether the content is doing some of the work.
A 2024 study comparing humans and language models found that a problem’s content affected performance in both across several reasoning tasks, although their errors differed. That doesn’t mean their machinery is identical. But if dependence on familiar content is evidence against reasoning, we have some explaining to do about people as well.
I like Douglas Hofstadter’s admission in “Is there an ‘I’ in AI?” that he judged a human relative’s flawed mathematical poem more generously than a smaller flaw from GPT-4. It’s easy to supply a sympathetic explanation for a person’s mistake. We assume they understand, so something must have gone wrong. With a model, the mistake can become evidence that it never understood anything. I can see why we make that distinction. I’d still like us to justify it.
Where did that explanation come from?
Graepel also wants explanations that faithfully describe how a conclusion was reached. I’d like that too. I’m just not sure being human entitles us to assume we can provide them.
In a choice-blindness experiment, people chose which of two faces they found more attractive. The researchers sometimes secretly handed them the photograph they hadn’t chosen, then asked them to explain their preference. Participants could give reasons for the substituted choice. They weren’t necessarily reporting the process that produced their original decision, even if they sincerely believed they were.
Models have a reporting problem too. In experiments that supplied hints, those hints could influence answers without being acknowledged in the reasoning text. These are different experiments, not evidence of a common mechanism. But in either case, asking for an explanation hasn’t settled whether the explanation is faithful.
If you give me a proof, I can check whether it works. That doesn’t tell me how you found it. Conversely, an honest account of how you found it wouldn’t make an invalid proof valid. We should be clear which of those things we’re asking a person—or a model—to give us.
If we want to know how a model reached its answer, we have to investigate the computation. Calling it “just next-token prediction” describes how it produces output, but leaves open what happens along the way. Formal work on transformers with chain of thought shows that intermediate generation can change computational power under specified assumptions. That is a capacity result, not proof that a trained model learned a particular algorithm.
One way to investigate that is to change something inside the model and see whether the answer changes. In a circuit-tracing study, answering a question about Dallas and its state’s capital involved an intermediate representation of Texas. Substituting California for Texas changed the answer to Sacramento. The method is imperfect and this is a selected example, but the intervention gives us evidence that the representation contributed to the answer. We can ask where that computation works, where it fails, and whether it survives changes to the problem.
Let me keep the notebook
There’s a part of Graepel’s proposal I find easier to agree with once I think about code review.
I’ve argued for programming languages that make code locally reviewable. Given a small change, how much else does the reviewer have to understand to judge its effect? If the answer is “most of the codebase,” asking for more conscientious reviewers won’t get us very far. The language and the tools should help make the change understandable.
The same question applies when I read an argument. A person can retain a belief, act on it, and revise it after learning something new. That doesn’t mean they can hand me an inspectable record of every belief and how it got there. But if they’ve written down the relevant evidence and steps, I can check the argument without that access.
Steven Pinker’s “Reason To Believe” makes the broader case for our dependence on expertise and institutions that expose claims to criticism. An awful lot of what we count as human knowledge depends on work spread across people, records, and tools. We ought to remember that when deciding where the human stops and the reasoning system begins.
So yes, give the machine explicit state and checked updates. Make it possible to inspect the evidence. These are useful engineering requirements even if we haven’t agreed that they define reasoning. Then evaluate the whole system, including whatever is supposed to catch its mistakes. A model approving its own explanation is a rather thin form of review.
“Humans make mistakes too” would be a terrible reason to trust a model with something consequential. Its competence still needs evidence. What I object to is treating a human’s mistakes as problems to investigate while treating a model’s mistakes as the end of the investigation.
I haven’t arrived at a definition of reasoning that I’m happy with. I’d like one that helps explain why some inferences work, what has to be learned, and what happens when the familiar cues disappear. If it excludes humans, or includes regex engines, let’s follow that through and see whether we still want it.
And if we’re testing my ability to reason, please let me use a pencil. We can argue about whether it’s part of the system afterwards.
I used Codex in this post to generate images and to help with fact-checking, grammar, and editing.




