4 papers
Measuring Explainer Stability via Attribution Separability
Eddie Conti, Álvaro Parafita, Axel Brando
Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scor…
CID: Measuring Feature Importance Through Counterfactual Distributions
Eddie Conti, Álvaro Parafita, Axel Brando
Assessing the importance of individual features in Machine Learning is critical to understand the model's decision-making process. While numerous methods exist, the lack of a defin…
Practical do-Shapley Explanations with Estimand-Agnostic Causal Inference
Álvaro Parafita, Tomas Garriga, Axel Brando +1
Among explainability techniques, SHAP stands out as one of the most popular, but often overlooks the causal structure of the problem. In response, do-SHAP employs interventional qu…
Probing the Embedding Space of Transformers via Minimal Token Perturbations
Eddie Conti, Alejandro Astruc, Alvaro Parafita +1
Understanding how information propagates through Transformer models is a key challenge for interpretability. In this work, we study the effects of minimal token perturbations on th…