← Writing

Do LLMs Know When Their Chain-of-Thought Is Wrong?

TL;DR — Yes — internally. A linear probe on hidden states predicts whether a chain-of-thought trace is correct with 0.95 AUROC (already 0.79 at the first reasoning step), while the model's verbalized confidence on wrong traces is 4.55/5 versus 4.87/5 on correct ones. A text-only classifier reaches just 0.59. But the signal is diagnostic, not causal: steering, probe-guided best-of-N, self-correction, and activation patching all fail to fix the errors.

Chain-of-thought prompting rests on an assumption: the reasoning a model writes out reflects the computation it actually performs. This paper shows that assumption breaks in a specific, measurable way — models know when their reasoning is going wrong, and don't say so.

The gap between what a model knows and what it says

The effect holds across three model families (Qwen, Llama, Phi), from 1.5B to 72B parameters, and in RL-trained reasoning models (DeepSeek-R1 at 0.852 AUROC).

Can the signal fix the errors? No.

If a model knows it is wrong, can we use that knowledge to make it right? The paper tests four interventions:

  1. Activation steering
  2. Probe-guided best-of-N
  3. Self-correction
  4. Activation patching — which destroys output coherence entirely

All four fail.

The signal is a readout of computation quality, not a lever to redirect it.

Why it matters

For reliability, the result is useful: a cheap probe can flag likely-wrong reasoning before you act on it — a natural input to a verification agent or a human-review gate. For interpretability, it draws a boundary: error representations during reasoning behave differently from the factual-knowledge representations that prior work has successfully edited.

Frequently asked questions

Do LLMs know when their reasoning is wrong?

Internally, largely yes. A linear probe on hidden states predicts chain-of-thought correctness with 0.95 AUROC, while the model's verbalized confidence stays high on wrong answers (4.55/5 vs 4.87/5 on correct ones).

Can hidden error signals be used to correct LLM reasoning?

Not in this study. Activation steering, probe-guided best-of-N, self-correction, and activation patching all failed; the signal is diagnostic, not causal.

Does hidden error awareness appear in reasoning models like DeepSeek-R1?

Yes. The effect holds across Qwen, Llama, and Phi from 1.5B to 72B parameters and in RL-trained reasoning models, with DeepSeek-R1 at 0.852 AUROC.


Based on Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal (arXiv 2026, arXiv:2605.09502) by Aojie Yuan, Zhiyuan Julian Su, Haiyue Zhang, Yi Nian, Yue Zhao. Written by Aojie (Justin) Yuan, USC Fortis Lab.