8 papers
Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages
Isabella Gidi, Antonio Almudévar, Core Francisco Park +2
Multilingual large language models often generalize across languages, and prior work suggests that their internal mechanisms can overlap cross-lingually. It remains unclear, howeve…
Do Activation Verbalization Methods Convey Privileged Information?
Millicent Li, Alberto Mario Ceballos Arroyo, Giordano Rogers +2
Recent interpretability methods have proposed to translate LLM internal representations into natural language descriptions using a second verbalizer LLM. This is intended to illumi…
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
Tian Qin, Naomi Saphra, David Alvarez-Melis
Early in training, LMs can behave like n-gram models, but eventually they often learn tree-based syntactic rules and generalize hierarchically out of distribution (OOD). We study t…
Can Interpretation Predict Behavior on Unseen Data?
Victoria R. Li, Jenny Kaufmann, Martin Wattenberg +3
Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this a…
PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
Oskar van der Wal, Pietro Lesci, Max Muller-Eberstein +4
The stability of language model pre-training and its effects on downstream performance are still understudied. Prior work shows that the training process can yield significantly di…
Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon
USVSN Sai Prashanth, Alvin Deng, Kyle O'Brien +9
Memorization in language models is typically treated as a homogenous phenomenon, neglecting the specifics of the memorized data. We instead model memorization as the effect of a se…