4 papers
Measuring Explainer Stability via Attribution Separability
Eddie Conti, Ãlvaro Parafita, Álvaro Parafita +1
Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scor…
CID: Measuring Feature Importance Through Counterfactual Distributions
Eddie Conti, Ãlvaro Parafita, Axel Brando
Assessing the importance of individual features in Machine Learning is critical to understand the model's decision-making process. While numerous methods exist, the lack of a defin…
Probing the Embedding Space of Transformers via Minimal Token Perturbations
Eddie Conti, Alejandro Astruc, Alvaro Parafita +1
Understanding how information propagates through Transformer models is a key challenge for interpretability. In this work, we study the effects of minimal token perturbations on th…
An alternative formulation of attention pooling function in translation
Eddie Conti
The aim of this paper is to present an alternative formulation of the attention scoring function in translation tasks. Generally speaking, language is deeply structured, and this i…