2 papers
cs.CV2022
Towards More Robust Interpretation via Local Gradient Alignment
Sunghwan Joo, Seokhyeon Jeong, Juyeon Heo +2
Neural network interpretation methods, particularly feature attribution methods, are known to be fragile with respect to adversarial input perturbations. To address this, several m…
cs.LG2019
Fooling Neural Network Interpretations via Adversarial Model Manipulation
Juyeon Heo, Sunghwan Joo, Taesup Moon
We ask whether the neural network interpretation methods can be fooled via adversarial model manipulation, which is defined as a model fine-tuning step that aims to radically alter…