← Writing

Is a Concept Subspace Just the Logit Lens in Disguise?

TL;DR — Not in the models tested. A concept subspace's effect on behavior doesn't tell you how it relates to the output readout, so this paper introduces a two-sided, weights-only geometric diagnostic: measure the subspace's overlap with the dominant right-singular directions of the unembedding matrix, against readout-oriented positive controls. Across nine estimators and 26 models, activation-derived concept estimators carry only 0.38–0.80% mean energy in the top-ten readout span, while readout-oriented controls carry far more.

Interpretability papers often extract a “concept subspace” and show that intervening on it changes model behavior. But that leaves a basic question open: is the subspace a genuinely upstream representation, or is it just aligned with the directions the model uses to emit output tokens — the logit lens in disguise?

A weights-only test

The diagnostic measures how much of an extracted subspace overlaps the dominant right-singular directions of the unembedding matrix (the top-ten readout span), using principal angles. Given an extracted basis, the raw diagnostic needs only model weights — no new forward passes.

It is two-sided: low overlap is only meaningful if readout-oriented subspaces register high overlap under the same test. So it is calibrated against two positive controls:

Results across 26 models

A complementary four-model, three-seed intervention study finds model-dependent source-directed effects that remain well below full-vector replacement.

Geometry locates the representation; interventions test behavioral effects — and neither alone licenses a claim of causal sufficiency.

This paper is the follow-up to Beyond Language, which introduced FARS.

Frequently asked questions

What is the logit lens?

A technique that reads intermediate hidden states through the model's unembedding matrix to see which output tokens they point toward.

How do you test whether a concept subspace is just aligned with the output readout?

Measure the principal angles between the subspace and the dominant right-singular directions of the unembedding matrix, and calibrate against readout-oriented positive controls such as final-layer PCA and a same-layer next-token fit. The raw test uses only model weights.

Are activation-derived concept subspaces aligned with the readout?

In this study, no: across 26 models they carried only 0.38–0.80% mean energy in the top-ten readout span, while a next-token control carried about thirteen times more energy than FARS.


Based on Concept Subspaces Compute Beyond the Logit Lens: A Weights-Only Test for Locating Representations Upstream of Readout (arXiv 2026, arXiv:2609.39263) by Aojie Yuan, Zhiyuan Julian Su, Haiyue Zhang, Zijian Su. Written by Aojie (Justin) Yuan, USC Fortis Lab.