5 papers
Validating Causal Abstraction Metrics on Simulated Complex Systems
Maxime Méloux, Tiago Pimentel, François Portet +1
A central goal of science is to produce valid explanations of complex systems: high-level causal accounts that faithfully reflect the behavior of lower-level mechanisms. Yet no con…
MIST: Mutual Information Estimation Via Supervised Training
German Gritsai, Megan Richards, Maxime Méloux +2
We propose a fully data-driven approach to designing mutual information (MI) estimators. Since any MI estimator is a function of the observed sample from two random variables, we p…
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
Maxime Méloux, François Portet, Maxime Peyrard
Mechanistic Interpretability (MI) aims to reverse-engineer model behaviors by identifying functional sub-networks. Yet, the scientific validity of these findings depends on their s…
The Dead Salmons of AI Interpretability
Maxime Méloux, Giada Dirupo, François Portet +1
In a striking neuroscience study, the authors placed a dead salmon in an MRI scanner and showed it images of humans in social situations. Astonishingly, standard analyses of the ti…
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
Maxime Méloux, Silviu Maniu, François Portet +1
As AI systems are used in high-stakes applications, ensuring interpretability is crucial. Mechanistic Interpretability (MI) aims to reverse-engineer neural networks by extracting h…