activity
20152024
most citedBuild It, Break It, Fix It: Contesting Secure Development

24 citations · 28 across the 5 of their papers we have counts for

collaborators

18 papers

cs.LG2024★ 1 cited

De-amplifying Bias from Differential Privacy in Language Model Fine-tuning

Sanjari Srivastava, Piotr Mardziel, Zhikhun Zhang +3

Fairness and privacy are two important values machine learning (ML) practitioners often seek to operationalize in models. Fairness aims to reduce model bias for social/demographic…

cs.CL2020

Influence Patterns for Explaining Information Flow in BERT

Kaiji Lu, Zifan Wang, Piotr Mardziel +1

While attention is all you need may be proving true, we do not know why: attention-based transformer models such as BERT are superior but how information flows from input tokens to…

cs.AI2020

Reconstructing Actions To Explain Deep Reinforcement Learning

Xuan Chen, Zifan Wang, Yucai Fan +4

Feature attribution has been a foundational building block for explaining the input feature importance in supervised learning with Deep Neural Network (DNNs), but face new challeng…

cs.IT2020

Fairness Under Feature Exemptions: Counterfactual and Observational Measures

Sanghamitra Dutta, Praveen Venkatesh, Piotr Mardziel +2

With the growing use of ML in highly consequential domains, quantifying disparity with respect to protected attributes, e.g., gender, race, etc., is important. While quantifying di…

cs.LG2020

Smoothed Geometry for Robust Attribution

Zifan Wang, Haofan Wang, Shakul Ramkumar +3

Feature attributions are a popular tool for explaining the behavior of Deep Neural Networks (DNNs), but have recently been shown to be vulnerable to attacks that produce divergent…

cs.CL2020★ 1 cited

Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language Models

Kaiji Lu, Piotr Mardziel, Klas Leino +2

LSTM-based recurrent neural networks are the state-of-the-art for many natural language processing (NLP) tasks. Despite their performance, it is unclear whether, or how, LSTMs lear…