activity
20122024
most citedFooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods

168 citations · 639 across the 30 of their papers we have counts for

collaborators
Showing 2022Show all

5 papers · 1 filter

cs.CL20227 cited

Successive Prompting for Decomposing Complex Questions

Dheeru Dua, Shivanshu Gupta, Sameer Singh +1

Answering complex questions that require making latent decisions is a challenging task, especially when limited supervision is available. Recent works leverage the capabilities of…

cs.LG20221 cited

Learning to Query Internet Text for Informing Reinforcement Learning Agents

Kolby Nottingham, Alekhya Pyla, Sameer Singh +1

Generalization to out of distribution tasks in reinforcement learning is a challenging problem. One successful approach improves generalization by conditioning policies on task or…

cs.CV20223 cited

ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension

Sanjay Subramanian, William Merrill, Trevor Darrell +3

Training a referring expression comprehension (ReC) model for a new visual domain requires collecting referring expressions, and potentially corresponding bounding boxes, for image…

cs.LG202228 cited

Rethinking Explainability as a Dialogue: A Practitioner's Perspective

Himabindu Lakkaraju, Dylan Slack, Yuxin Chen +2

As practitioners increasingly deploy machine learning models in critical domains such as health care, finance, and policy, it becomes vital to ensure that domain experts function e…

cs.CL20225 cited

Identifying Adversarial Attacks on Text Classifiers

Zhouhang Xie, Jonathan Brophy, Adam Noack +6

The landscape of adversarial attacks against text classifiers continues to grow, with new attacks developed every year and many of them available in standard toolkits, such as Text…