3 papers
cs.LG2026
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
Moritz Miller, Florent Draye, Bernhard Schölkopf +1
A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to suppor…
cs.CL2026
On the Emergence and Test-Time Use of Structural Information in Large Language Models
Michelle Chao Chen, Moritz Miller, Bernhard Schölkopf +1
Learning structural information from observational data is central to producing new knowledge outside the training corpus. This holds for mechanistic understanding in scientific di…
cs.CL2025
Counterfactual reasoning: an analysis of in-context emergence
Moritz Miller, Bernhard Schölkopf, Siyuan Guo
Large-scale neural language models exhibit remarkable performance in in-context learning: the ability to learn and reason about the input context on the fly. This work studies in-c…