4 papers
Debiasing Without Protected Attributes: Latent Concept Erasure from Textual Profiles
Shun Shao, Zheng Zhao, Anna Korhonen +2
Most fairness research in NLP assumes direct access to protected attributes such as gender, race, or nationality. In practice, however, such information is often unavailable due to…
Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer
Shun Shao, Binxu Wang, Shay B. Cohen +2
Mechanistic interpretability has made it possible to localize circuits underlying specific behaviors in language models, but existing methods are expensive, model-specific, and dif…
Iterative Multilingual Spectral Attribute Erasure
Shun Shao, Yftah Ziser, Zheng Zhao +3
Multilingual representations embed words with similar meanings to share a common semantic space across languages, creating opportunities to transfer debiasing effects between langu…
MIB: A Mechanistic Interpretability Benchmark
Aaron Mueller, Atticus Geiger, Sarah Wiegreffe +20
How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretabili…