1 citations · 2 across the 2 of their papers we have counts for
3 papers
cs.LG2025★ 1 cited
MIB: A Mechanistic Interpretability Benchmark
Aaron Mueller, Atticus Geiger, Sarah Wiegreffe +20
How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretabili…
cs.CL2025
Are formal and functional linguistic mechanisms dissociated in language models?
Michael Hanna, Yonatan Belinkov, Sandro Pezzelle
Although large language models (LLMs) are increasingly capable, these capabilities are unevenly distributed: they excel at formal linguistic tasks, such as producing fluent, gramma…
cs.CL2024★ 1 cited
Incremental Sentence Processing Mechanisms in Autoregressive Transformer Language Models
Michael Hanna, Aaron Mueller
Autoregressive transformer language models (LMs) possess strong syntactic abilities, often successfully handling phenomena from agreement to NPI licensing. However, the features th…