Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Attention Mean Fields Predict Average Representation Dynamics and Reveal Context-Specific Computation
Micah Adler, John W. Byers, Mark Crovella
A language model's representation geometry is not predetermined; it evolves as the model runs. A faithful account of that geometry must capture that dynamic process, and so cannot…
cs.LG2026
Singular Vectors of Attention Heads Align with Features
Gabriel Franco, Carson Loughridge, Mark Crovella
Identifying feature representations in language models is a central task in mechanistic interpretability. Several recent studies have made the observation that feature representati…
cs.LG2026
Finding Interpretable Prompt-Specific Circuits in Language Models
Gabriel Franco, Lucas M. Tassis, Azalea Rohr +1
Understanding the internal circuits that language models use to solve tasks remains a central challenge in mechanistic interpretability. A crucial part of finding circuits is under…