activity
20242026
collaborators

8 papers

cs.CL2026

Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages

Isabella Gidi, Antonio Almudévar, Core Francisco Park +2

Multilingual large language models often generalize across languages, and prior work suggests that their internal mechanisms can overlap cross-lingually. It remains unclear, howeve…

cs.CL2026

Do Activation Verbalization Methods Convey Privileged Information?

Millicent Li, Alberto Mario Ceballos Arroyo, Giordano Rogers +2

Recent interpretability methods have proposed to translate LLM internal representations into natural language descriptions using a second verbalizer LLM. This is intended to illumi…

cs.LG2025

Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization

Tian Qin, Naomi Saphra, David Alvarez-Melis

Early in training, LMs can behave like n-gram models, but eventually they often learn tree-based syntactic rules and generalize hierarchically out of distribution (OOD). We study t…

cs.LG2025

Can Interpretation Predict Behavior on Unseen Data?

Victoria R. Li, Jenny Kaufmann, Martin Wattenberg +3

Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this a…

cs.CL2025

PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs

Oskar van der Wal, Pietro Lesci, Max Muller-Eberstein +4

The stability of language model pre-training and its effects on downstream performance are still understudied. Prior work shows that the training process can yield significantly di…

cs.CL2025

Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon

USVSN Sai Prashanth, Alvin Deng, Kyle O'Brien +9

Memorization in language models is typically treated as a homogenous phenomenon, neglecting the specifics of the memorized data. We instead model memorization as the effect of a se…