15 citations · 20 across the 2 of their papers we have counts for
6 papers
ML-LOO: Detecting Adversarial Examples with Feature Attribution
Puyudi Yang, Jianbo Chen, Cho-Jui Hsieh +2
Deep neural networks obtain state-of-the-art performance on a series of tasks. However, they are easily fooled by adding a small adversarial perturbation to input. The perturbation…
HopSkipJumpAttack: A Query-Efficient Decision-Based Attack
Jianbo Chen, Michael I. Jordan, Martin J. Wainwright
The goal of a decision-based adversarial attack on a trained model is to generate adversarial examples based solely on observing output labels returned by the targeted model. We de…
LS-Tree: Model Interpretation When the Data Are Linguistic
Jianbo Chen, Michael I. Jordan
We study the problem of interpreting trained classification models in the setting of linguistic data sets. Leveraging a parse tree, we propose to assign least-squares based importa…
L-Shapley and C-Shapley: Efficient Model Interpretation for Structured Data
Jianbo Chen, Le Song, Martin J. Wainwright +1
We study instancewise feature importance scoring as a method for model interpretation. Any such method yields, for each predicted instance, a vector of importance scores associated…
Greedy Attack and Gumbel Attack: Generating Adversarial Examples for Discrete Data
Puyudi Yang, Jianbo Chen, Cho-Jui Hsieh +2
We present a probabilistic framework for studying adversarial attacks on discrete data. Based on this framework, we derive a perturbation-based method, Greedy Attack, and a scalabl…
Learning to Explain: An Information-Theoretic Perspective on Model Interpretation
Jianbo Chen, Le Song, Martin J. Wainwright +1
We introduce instancewise feature selection as a methodology for model interpretation. Our method is based on learning a function to extract a subset of features that are most info…