24 citations · 26 across the 3 of their papers we have counts for
10 papers
Reconstructing Actions To Explain Deep Reinforcement Learning
Xuan Chen, Zifan Wang, Yucai Fan +4
Feature attribution has been a foundational building block for explaining the input feature importance in supervised learning with Deep Neural Network (DNNs), but face new challeng…
Smoothed Geometry for Robust Attribution
Zifan Wang, Haofan Wang, Shakul Ramkumar +3
Feature attributions are a popular tool for explaining the behavior of Deep Neural Networks (DNNs), but have recently been shown to be vulnerable to attacks that produce divergent…
Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language Models
Kaiji Lu, Piotr Mardziel, Klas Leino +2
LSTM-based recurrent neural networks are the state-of-the-art for many natural language processing (NLP) tasks. Despite their performance, it is unclear whether, or how, LSTMs lear…
Interpreting Interpretations: Organizing Attribution Methods by Criteria
Zifan Wang, Piotr Mardziel, Anupam Datta +1
Motivated by distinct, though related, criteria, a growing number of attribution methods have been developed tointerprete deep learning. While each relies on the interpretability o…
Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks
Haofan Wang, Zifan Wang, Mengnan Du +5
Recently, increasing attention has been drawn to the internal mechanisms of convolutional neural networks, and the reason why the network makes specific decisions. In this paper, w…
Build It, Break It, Fix It: Contesting Secure Development
James Parker, Michael Hicks, Andrew Ruef +5
Typical security contests focus on breaking or mitigating the impact of buggy systems. We present the Build-it, Break-it, Fix-it (BIBIFI) contest, which aims to assess the ability…