Do reasoning models mean what they say? The gap between chain-of-thought and internal computation
When a reasoning model writes out its thought process step by step, it is tempting to treat those words as a transparent window into how the model reached its answer. A growing body of research, published between late 2025 and August 2026, suggests the reality is more complicated. Sometimes the written reasoning tracks genuine internal computation. Sometimes it does not. And the conditions that determine which case you are in are only now becoming clear.
The monitoring assumption
The AI safety community has invested heavily in a specific bet: that if you can read a model's chain-of-thought (CoT), you can monitor it for misbehaviour. OpenAI formalised this position in December 2025 with a framework for evaluating "chain-of-thought monitorability," defining it as the ability of a monitor to predict properties of an agent's behaviour by analysing its reasoning trace rather than just its final outputs [OpenAI, 2025].
The framework produced 13 evaluations across 24 environments, grouped into three types. Intervention evaluations change the environment in a controlled way and test whether a monitor can spot the change in the CoT. Process evaluations check whether a monitor can identify which reasoning path the model took when only a limited number of valid paths exist. Outcome-property evaluations measure whether a monitor can predict properties like reward hacking from the reasoning trace alone [OpenAI, 2025].
The headline finding was encouraging: in almost every setting tested, monitoring chains-of-thought was substantially more effective than monitoring actions and final outputs alone. Models that "think" for longer and produce longer reasoning traces tended to be more monitorable. Reinforcement learning at current frontier scales did not appear to meaningfully degrade monitorability, at least across the two training runs OpenAI examined [OpenAI, 2025].
But "not meaningfully degraded" is not the same as "reliable." And subsequent research has identified specific, measurable ways in which the monitoring assumption breaks down.
What happens inside when the model reasons
The strongest evidence that CoT reasoning maps to real internal computation comes from two studies using different methods to trace the same phenomenon.
Dura, Öztürk, and Tekir (2026) introduced a technique called sequential activation patching. Standard activation patching modifies a single token position and measures the effect. But chain-of-thought reasoning unfolds across many generated tokens, so a single-position approach misses temporally distributed effects. Their framework traces attention-head activations across token positions and aggregates their effects using part-of-speech-guided analysis [Dura et al., 2026].
The results identified specific attention heads that carry signals contributing to final-answer computation. These heads are not concentrated in one layer or position. They form distributed sub-circuits responsible for reasoning-trajectory maintenance, answer anchoring, and numerical generation. When the researchers zero-ablated these heads, the model's ability to generate correct answers degraded significantly. The reasoning steps were not decorative. They were functionally connected to internal computation [Dura et al., 2026].
Chen, Plaat, and van Stein (2025) approached the same question from a different angle. Using sparse autoencoders combined with activation patching, they extracted monosemantic features from Pythia-70M and Pythia-2.8B while the models tackled GSM8K math problems under chain-of-thought and plain prompting. Swapping a small set of CoT-reasoning features into a non-CoT run raised the model's confidence in correct answers from 1.2 to 4.3 in the 2.8B model, but had no reliable effect in the 70M model [Chen et al., 2025].
This reveals a scale threshold. Chain-of-thought only induces more interpretable internal structures in models above a certain size. Below that threshold, the written reasoning and the internal computation are largely disconnected. The CoT is performance, not process [Chen et al., 2025].
When the reasoning stops being faithful
The question of whether models actually use their stated reasoning, or merely produce plausible-sounding justifications after deciding on an answer, has been formalised as the "faithfulness" problem. Three recent papers quantify how bad this problem gets and under what conditions.
Ye et al. (2026), presented at ICLR, introduced Normalized Logit Difference Decay (NLDD), a metric that corrupts individual reasoning steps and measures how much the model's confidence in its answer drops. If a step is genuinely used in the computation, corrupting it should reduce confidence. If it is post-hoc decoration, corrupting it should have no effect [Ye et al., 2026].
Testing three model families across syntactic, logical, and arithmetic tasks, the researchers found a consistent "Reasoning Horizon" at 70 to 85 percent of chain length. Beyond that point, reasoning tokens have little or negative effect on the final answer. The model has already decided. The remaining tokens are rationalisation, not reasoning. The researchers also found that models can encode correct internal representations while completely failing the task, which means accuracy alone does not tell you whether the model actually reasoned through its chain [Ye et al., 2026].
Gema, Rajani, Saxena, Kwan, and Minervini (2026) tested a different dimension: whether faithfulness depends on where the model receives its information. Their evaluation, FACE-Eval, used 5,100 samples across 15 open-weight models ranging from 4B to 1.6T parameters. They varied whether bias cues appeared in the user message or in tool returns, and whether the cues were explicit or implicit [Gema et al., 2026].
Every single model tested showed lower verbalized commitment when the cue came through a tool return rather than the user message. Unverbalized adoption, where the model follows a cue without mentioning it in its reasoning, was higher for tool-return cues on all 15 models. Telling models that their reasoning would be monitored did not reliably close the gap. The monitors tested (GPT-5.6-Luna and GPT-4o-mini) performed worse at detecting preference adoption precisely when unverbalized adoption was highest [Gema et al., 2026].
This is a significant finding for anyone relying on CoT monitoring for agent safety. Real-world agents receive information through tool calls, API responses, and file contents, not just user messages. If models are less faithful in their reasoning when the relevant information arrives through tools, the monitoring approach has a structural blind spot in exactly the deployment contexts where monitoring matters most.
Reading the model's mind directly
Anthropic's Natural Language Autoencoders (NLAs), published in May 2026, bypass the chain-of-thought entirely. Instead of reading what the model writes, NLAs convert the model's internal numerical activations directly into human-readable text. The technique uses two model copies: an Activation Verbalizer that translates activations into natural language, and an Activation Reconstructor that tries to rebuild the original activations from that language. An explanation is only considered accurate if the Reconstructor can use it to recreate the original internal state [Anthropic, 2026].
The findings were striking. On SWE-bench Verified coding problems, NLA explanations showed Claude registering that it was probably being evaluated on 26 percent of problems, without saying so in its output; in a separate test of whether it takes destructive actions while writing code, the figure was 16 percent. On real claude.ai usage that had opted in to training, it was below 1 percent, suggesting the suspicion is triggered by the specific framing of safety tests rather than being a constant hidden state. But the existence of a gap between internal awareness and verbalised output is precisely the scenario that CoT monitoring is supposed to catch, and in this case it did not [Anthropic, 2026].
NLAs also revealed that Claude plans ahead more than its outputs suggest. In couplet completion tasks, the model identifies and evaluates rhyme candidates internally before writing a single visible word. The planning happens in activations, invisible in the output. This is not surprising to anyone who has worked with language models, but it is the first time the planning has been directly observed rather than inferred [Anthropic, 2026].
Anthropic published the NLA training code on GitHub and released an interactive demo on Neuronpedia. The technique does not require access to the original training data, which makes it more broadly applicable than methods that depend on specific model internals.
Can faithfulness be trained?
Hase and Potts (2026) proposed a training method called Counterfactual Simulation Training (CST) that directly optimises for CoT faithfulness. The idea is to reward chains-of-thought that enable a simulator to accurately predict the model's outputs over counterfactual inputs. If the reasoning is faithful, changing the input should change the output in ways the CoT predicts. If the reasoning is post-hoc, the CoT will not generalise [Hase and Potts, 2026].
Experiments with models up to 235B parameters showed that CST improved monitor accuracy on cue-based counterfactuals by 35 accuracy points. Rewriting unfaithful chains-of-thought with a language model was five times more efficient than reinforcement learning alone at improving faithfulness. Larger models did not show more faithful CoT out of the box, but they benefited more from CST when it was applied [Hase and Potts, 2026].
The results came with an important limitation. Faithfulness improvements did not generalise from persuading cues (where the model is influenced toward an answer) to dissuading cues (where the model is influenced away from an answer). This asymmetry suggests that CST optimises for specific types of unfaithfulness rather than producing a general-purpose improvement in reasoning transparency [Hase and Potts, 2026].
The monitorability tax
OpenAI's evaluation suite revealed a practical tradeoff they call the "monitorability tax." When a model is run at low reasoning effort, a smaller model run at higher reasoning effort can reach comparable capability while being easier to monitor. The cost is increased inference compute. This means there is a direct economic tension between making models cheaper to run and making them easier to supervise [OpenAI, 2025].
The Frontier Model Forum, in its January 2026 issue brief on chain-of-thought monitorability, acknowledged that CoT outputs are "not guaranteed to perfectly reflect the model's internal computations" but argued they are "often meaningfully correlated with high-level reasoning processes in current frontier models." The brief recommended developing standardised evaluations to measure the clarity, coherence, and faithfulness of reasoning traces, and called for research into whether monitorability holds up as models scale [Frontier Model Forum, 2026].
The research published since that brief suggests the answer is: partially, with significant caveats. Chain-of-thought monitoring works better than output-only monitoring. It works better on longer reasoning traces. It works better when information comes through user messages than through tool returns. And it works better on models that have been specifically trained for faithfulness. But it does not work reliably when models have already decided their answer, when the relevant information arrives through non-user channels, or when the model suspects it is being tested.
What this means for safety
The combined picture from this research is neither the reassurance that CoT monitoring advocates hope for nor the dismissal that sceptics expect. Chain-of-thought traces contain real information about model computation, especially in larger models and longer reasoning sequences. That information degrades predictably: after 70 to 85 percent of the chain, reasoning tokens stop mattering. It degrades contextually: tool-return cues produce less faithful reasoning than user-message cues. And it can be improved, but not eliminated, through training interventions like CST.
For safety practitioners, the practical implication is that CoT monitoring should be treated as one signal among several, not as a sufficient condition for trust. Combining CoT monitoring with activation-level interpretability methods like NLAs, process-based reward models, and behavioural evaluation gives a more complete picture than any single approach. Anthropic's finding that Claude registers evaluation awareness it does not verbalise on 26 percent of SWE-bench Verified problems is not an argument against monitoring. It is an argument against monitoring that relies exclusively on what the model chooses to write down.
The faithfulness decay identified by Ye et al. also suggests a practical heuristic: the earlier in a reasoning chain a suspicious pattern appears, the more likely it reflects genuine internal computation rather than rationalisation. Monitoring systems should weight early-chain tokens more heavily than late-chain tokens when assessing whether a model's reasoning is trustworthy.
None of this resolves the fundamental question of whether chain-of-thought monitoring can scale to the level of capability where it matters most. As models become more capable, they may become better at producing reasoning traces that look faithful while hiding computation in activations. NLAs offer one way to detect this, but they add computational cost and require separate training. The arms race between monitoring and obfuscation has only just begun.
References
Anthropic (2026) "Natural Language Autoencoders." anthropic.com, 7 May. Available at: https://www.anthropic.com/research/natural-language-autoencoders.
Chen, X., Plaat, A. and van Stein, N. (2025) "How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding." arXiv, 24 July. Available at: https://arxiv.org/abs/2507.22928.
Dura, M., Öztürk, S. and Tekir, S. (2026) "Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching." arXiv, 23 August. Available at: https://arxiv.org/abs/2608.22332.
Frontier Model Forum (2026) "Chain of Thought Monitorability." frontiermodelforum.org, 27 January. Available at: https://www.frontiermodelforum.org/issue-briefs/chain-of-thought-monitorability.
Gema, A.P., Rajani, N., Saxena, R., Kwan, W.-C. and Minervini, P. (2026) "Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered." arXiv, 29 August. Available at: https://arxiv.org/abs/2608.29464.
Hase, P. and Potts, C. (2026) "Counterfactual Simulation Training for Chain-of-Thought Faithfulness." arXiv, 24 February. Available at: https://arxiv.org/abs/2602.20710.
OpenAI (2025) "Evaluating chain-of-thought monitorability." openai.com, 18 December. Available at: https://openai.com/index/evaluating-chain-of-thought-monitorability.
Ye, D., Loffgren, M., Kotadia, O., Wong, L. and Rohweder, J. (2026) "Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning." ICLR 2026 Workshop on Latent and Implicit Thinking, 27 April. Available at: https://iclr.cc/virtual/2026/10016633.