4 papers
Lagged Coupling: Internal Representations Become Readable Before They Become Causal
Xining Xun
Across the full Pythia suite (160M-12B, eight checkpoints, four task families), a linear probe can read a target variable from the residual stream as early as step 1,000 at every s…
Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release
Xining Xun
A growing body of work reports that language models represent task-relevant latent structure that they fail to use. Whether such structure, once located, can be converted into beha…
Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library
Xining Xun
When a language model answers an interventional question, the computation it must perform depends on the type of evidence the query requires. We report a decoupling in how a transf…
Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
Xining Xun
Interventional data is widely regarded as the gold standard for teaching models causal reasoning. We test this assumption in a fully controlled synthetic environment pitting observ…