1 paper · 1 filter
Naomi Saphra, Sarah Wiegreffe
The rise of the term "mechanistic interpretability" has accompanied increasing interest in understanding neural models -- particularly language models. However, this jargon has als…