← Writing

Do LLMs Reason the Same Way in English, Code, and Math?

TL;DR — Largely, yes. Across five LLMs (1.6B–8B), the same reasoning concept written as English prose, Python code, or math notation lands in a shared Format-Agnostic Reasoning Subspace (FARS) in the middle layers. A 10-dimensional subspace amplifies concept structure 3× while suppressing form; replacing only those 10 dimensions during cross-form patching preserves 90–96% of model output, versus 44–56% for full activation replacement. The main split is declarative vs procedural, not language vs formal.

“If the switch is on, the lamp lights” can be written as an English sentence, a line of Python, or a formula. Does a language model represent these as one idea or three?

The TriForm benchmark

The paper introduces TriForm: 18 concepts × 6 forms × 3 instances = 324 stimuli, and studies five LLMs (1.6B–8B) from three architecture families using permutation-corrected RSA, cross-form probing, and activation patching.

A 10-dimensional shared subspace

All three methods converge on a Format-Agnostic Reasoning Subspace (FARS) in the middle layers. Concept-centroid PCA extracts it as a 10-dimensional basis that amplifies concept structure about 3× while pushing form information to near zero.

Declarative vs procedural

The critical axis of divergence is not linguistic vs formal, but declarative vs procedural.

Representations are far more compatible between prose and mathematics than between either and code. The follow-up paper, Concept Subspaces Compute Beyond the Logit Lens, tests where FARS sits relative to the model's output readout.

Frequently asked questions

What is a Format-Agnostic Reasoning Subspace (FARS)?

A 10-dimensional subspace in an LLM's middle layers, extracted by concept-centroid PCA, that encodes a reasoning concept the same way whether it is written as prose, code, or mathematical notation.

Do LLMs share representations across natural language, code, and math?

Largely yes: in five LLMs, replacing only the 10 FARS dimensions during cross-form patching preserved 90–96% of model output, far more than full activation replacement (44–56%).

Which formats are most different inside LLMs?

Code. Prose and math are much more compatible with each other than either is with code, suggesting the key split is declarative vs procedural.


Based on Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models (arXiv 2026, arXiv:2605.09496) by Aojie Yuan, Zhiyuan Su. Written by Aojie (Justin) Yuan, USC Fortis Lab.