7 papers
Functions on Irreducible Components of the Emerton-Gee Stack
Eivind Otto Hjelle, Louis Jaburi, Rachel Knak +2
Let be a finite unramified extension, and let denote the Emerton-Gee stack parametrizing étale -modules of rank . It is known sinc…
Capability Provenance in Language Models: A Case Study in Social Reasoning
Glenn Matlin, Chandreyi Chakraborty, Saehee Eom +8
We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning i…
Bergson: An Open Source Library for Data Attribution
Lucia Quirke, Louis Jaburi, David Johnston +6
Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging unde…
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
Matthew Kowal, Goncalo Paulo, Louis Jaburi +6
As large language models are increasingly trained and fine-tuned, practitioners need methods to identify which training data drive specific behaviors, particularly unintended ones.…
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
Kola Ayonrinde, Louis Jaburi
Mechanistic Interpretability (MI) aims to understand neural networks through causal explanations. Though MI has many explanation-generating methods, progress has been limited by th…
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
Kola Ayonrinde, Louis Jaburi
Mechanistic Interpretability aims to understand neural networks through causal explanations. We argue for the Explanatory View Hypothesis: that Mechanistic Interpretability researc…