Do LLMs Reason the Same Way in English, Code, and Math?
“If the switch is on, the lamp lights” can be written as an English sentence, a line of Python, or a formula. Does a language model represent these as one idea or three?
The TriForm benchmark
The paper introduces TriForm: 18 concepts × 6 forms × 3 instances = 324 stimuli, and studies five LLMs (1.6B–8B) from three architecture families using permutation-corrected RSA, cross-form probing, and activation patching.
A 10-dimensional shared subspace
All three methods converge on a Format-Agnostic Reasoning Subspace (FARS) in the middle layers. Concept-centroid PCA extracts it as a 10-dimensional basis that amplifies concept structure about 3× while pushing form information to near zero.
- Replacing only those 10 dimensions in cross-form patching preserves 90–96% of model output.
- Full activation replacement preserves only 44–56%; variance-maximizing PCA, 60–74%.
- Ablating the subspace causes targeted disruption.
- FARS generalizes to held-out concepts and converges across architectures (CCA > 0.79 for all model pairs) — within-modality evidence for the Platonic Representation Hypothesis.
Declarative vs procedural
The critical axis of divergence is not linguistic vs formal, but declarative vs procedural.
Representations are far more compatible between prose and mathematics than between either and code. The follow-up paper, Concept Subspaces Compute Beyond the Logit Lens, tests where FARS sits relative to the model's output readout.
Frequently asked questions
What is a Format-Agnostic Reasoning Subspace (FARS)?
A 10-dimensional subspace in an LLM's middle layers, extracted by concept-centroid PCA, that encodes a reasoning concept the same way whether it is written as prose, code, or mathematical notation.
Do LLMs share representations across natural language, code, and math?
Largely yes: in five LLMs, replacing only the 10 FARS dimensions during cross-form patching preserved 90–96% of model output, far more than full activation replacement (44–56%).
Which formats are most different inside LLMs?
Code. Prose and math are much more compatible with each other than either is with code, suggesting the key split is declarative vs procedural.
Based on Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models (arXiv 2026, arXiv:2605.09496) by Aojie Yuan, Zhiyuan Su. Written by Aojie (Justin) Yuan, USC Fortis Lab.