most citedML-LOO: Detecting Adversarial Examples with Feature Attribution

15 citations · 20 across the 2 of their papers we have counts for

collaborators

6 papers

cs.LG201915 cited

ML-LOO: Detecting Adversarial Examples with Feature Attribution

Puyudi Yang, Jianbo Chen, Cho-Jui Hsieh +2

Deep neural networks obtain state-of-the-art performance on a series of tasks. However, they are easily fooled by adding a small adversarial perturbation to input. The perturbation…

cs.LG2019

HopSkipJumpAttack: A Query-Efficient Decision-Based Attack

Jianbo Chen, Michael I. Jordan, Martin J. Wainwright

The goal of a decision-based adversarial attack on a trained model is to generate adversarial examples based solely on observing output labels returned by the targeted model. We de…

cs.LG20195 cited

LS-Tree: Model Interpretation When the Data Are Linguistic

Jianbo Chen, Michael I. Jordan

We study the problem of interpreting trained classification models in the setting of linguistic data sets. Leveraging a parse tree, we propose to assign least-squares based importa…

cs.LG2018

L-Shapley and C-Shapley: Efficient Model Interpretation for Structured Data

Jianbo Chen, Le Song, Martin J. Wainwright +1

We study instancewise feature importance scoring as a method for model interpretation. Any such method yields, for each predicted instance, a vector of importance scores associated…

cs.LG2018

Greedy Attack and Gumbel Attack: Generating Adversarial Examples for Discrete Data

Puyudi Yang, Jianbo Chen, Cho-Jui Hsieh +2

We present a probabilistic framework for studying adversarial attacks on discrete data. Based on this framework, we derive a perturbation-based method, Greedy Attack, and a scalabl…

cs.LG2018

Learning to Explain: An Information-Theoretic Perspective on Model Interpretation

Jianbo Chen, Le Song, Martin J. Wainwright +1

We introduce instancewise feature selection as a methodology for model interpretation. Our method is based on learning a function to extract a subset of features that are most info…