2 papers
cs.LG2019
Robust Attribution Regularization
Jiefeng Chen, Xi Wu, Vaibhav Rastogi +2
An emerging problem in trustworthy machine learning is to train models that produce robust interpretations for their predictions. We take a step towards solving this problem throug…
cs.LG2018
Concise Explanations of Neural Networks using Adversarial Training
Prasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury +2
We show new connections between adversarial learning and explainability for deep neural networks (DNNs). One form of explanation of the output of a neural network model in terms of…