4 papers
Singular Vectors of Attention Heads Align with Features
Gabriel Franco, Carson Loughridge, Mark Crovella
Identifying feature representations in language models is a central task in mechanistic interpretability. Several recent studies have made the observation that feature representati…
Finding Interpretable Prompt-Specific Circuits in Language Models
Gabriel Franco, Lucas M. Tassis, Azalea Rohr +1
Understanding the internal circuits that language models use to solve tasks remains a central challenge in mechanistic interpretability. A crucial part of finding circuits is under…
Testing the Limits of Truth Directions in LLMs
Angelos Poulis, Mark Crovella, Evimaria Terzi
Large language models (LLMs) have been shown to encode truth of statements in their activation space along a linear truth direction. Previous studies have argued that these directi…
The workflow motif: a widely-useful performance diagnosis abstraction for distributed applications
Mania Abdi, Peter Desnoyers, Mark Crovella +1
Diagnosing problems in deployed distributed applications continues to grow more challenging. A significant reason is the extreme mismatch between the powerful abstractions develope…