Do LLMs Know When Their Chain-of-Thought Is Wrong?
Chain-of-thought prompting rests on an assumption: the reasoning a model writes out reflects the computation it actually performs. This paper shows that assumption breaks in a specific, measurable way — models know when their reasoning is going wrong, and don't say so.
The gap between what a model knows and what it says
- A linear probe on hidden states predicts trace correctness at 0.95 AUROC — and at 0.79 from the very first reasoning step.
- Verbalized confidence barely moves: 4.55/5 on wrong traces vs 4.87/5 on correct ones.
- A classifier on the generated text alone reaches only 0.59 — the error signal is largely invisible in the output.
The effect holds across three model families (Qwen, Llama, Phi), from 1.5B to 72B parameters, and in RL-trained reasoning models (DeepSeek-R1 at 0.852 AUROC).
Can the signal fix the errors? No.
If a model knows it is wrong, can we use that knowledge to make it right? The paper tests four interventions:
- Activation steering
- Probe-guided best-of-N
- Self-correction
- Activation patching — which destroys output coherence entirely
All four fail.
The signal is a readout of computation quality, not a lever to redirect it.
Why it matters
For reliability, the result is useful: a cheap probe can flag likely-wrong reasoning before you act on it — a natural input to a verification agent or a human-review gate. For interpretability, it draws a boundary: error representations during reasoning behave differently from the factual-knowledge representations that prior work has successfully edited.
Frequently asked questions
Do LLMs know when their reasoning is wrong?
Internally, largely yes. A linear probe on hidden states predicts chain-of-thought correctness with 0.95 AUROC, while the model's verbalized confidence stays high on wrong answers (4.55/5 vs 4.87/5 on correct ones).
Can hidden error signals be used to correct LLM reasoning?
Not in this study. Activation steering, probe-guided best-of-N, self-correction, and activation patching all failed; the signal is diagnostic, not causal.
Does hidden error awareness appear in reasoning models like DeepSeek-R1?
Yes. The effect holds across Qwen, Llama, and Phi from 1.5B to 72B parameters and in RL-trained reasoning models, with DeepSeek-R1 at 0.852 AUROC.
Based on Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal (arXiv 2026, arXiv:2605.09502) by Aojie Yuan, Zhiyuan Julian Su, Haiyue Zhang, Yi Nian, Yue Zhao. Written by Aojie (Justin) Yuan, USC Fortis Lab.