Is a Concept Subspace Just the Logit Lens in Disguise?
Interpretability papers often extract a “concept subspace” and show that intervening on it changes model behavior. But that leaves a basic question open: is the subspace a genuinely upstream representation, or is it just aligned with the directions the model uses to emit output tokens — the logit lens in disguise?
A weights-only test
The diagnostic measures how much of an extracted subspace overlaps the dominant right-singular directions of the unembedding matrix (the top-ten readout span), using principal angles. Given an extracted basis, the raw diagnostic needs only model weights — no new forward passes.
It is two-sided: low overlap is only meaningful if readout-oriented subspaces register high overlap under the same test. So it is calibrated against two positive controls:
- Final-layer PCA
- A same-layer next-token control, evaluated with a fitted linear translator for depth matching
Results across 26 models
- Four activation-derived concept estimators carry only 0.38–0.80% mean energy in the top-ten readout span (nine rank-matched estimators in total).
- Final-layer PCA carries 3.56%, exceeding the FARS subspace in 25 of 26 models.
- The next-token control carries roughly thirteen times more energy than FARS, with separation in all 25 tested models.
- Re-extracting FARS on ten disjoint concepts yields 62–100% cross-format retrieval across 24 generative models — the procedure transfers, not a fixed basis.
A complementary four-model, three-seed intervention study finds model-dependent source-directed effects that remain well below full-vector replacement.
Geometry locates the representation; interventions test behavioral effects — and neither alone licenses a claim of causal sufficiency.
This paper is the follow-up to Beyond Language, which introduced FARS.
Frequently asked questions
What is the logit lens?
A technique that reads intermediate hidden states through the model's unembedding matrix to see which output tokens they point toward.
How do you test whether a concept subspace is just aligned with the output readout?
Measure the principal angles between the subspace and the dominant right-singular directions of the unembedding matrix, and calibrate against readout-oriented positive controls such as final-layer PCA and a same-layer next-token fit. The raw test uses only model weights.
Are activation-derived concept subspaces aligned with the readout?
In this study, no: across 26 models they carried only 0.38–0.80% mean energy in the top-ten readout span, while a next-token control carried about thirteen times more energy than FARS.
Based on Concept Subspaces Compute Beyond the Logit Lens: A Weights-Only Test for Locating Representations Upstream of Readout (arXiv 2026, arXiv:2609.39263) by Aojie Yuan, Zhiyuan Julian Su, Haiyue Zhang, Zijian Su. Written by Aojie (Justin) Yuan, USC Fortis Lab.