4 papers
Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance
Sinie van der Ben, Neele Roch, Anna Hedström +1
Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each fe…
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
Anna Hedström, Salim I. Amoukou, Tom Bewley +2
We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventi…
Benchmarking XAI Explanations with Human-Aligned Evaluations
Rémi Kazmierczak, Steve Azzolin, Eloïse Berthier +9
We introduce PASTA (Perceptual Assessment System for explanaTion of Artificial Intelligence), a novel human-centric framework for evaluating eXplainable AI (XAI) techniques in comp…
From Flexibility to Manipulation: The Slippery Slope of XAI Evaluation
Kristoffer Wickstrøm, Marina Marie-Claire Höhne, Anna Hedström
The lack of ground truth explanation labels is a fundamental challenge for quantitative evaluation in explainable artificial intelligence (XAI). This challenge becomes especially p…