activity
20162022
most citedExSum: From Local Explanations to Model Understanding

1 citations · 1 across the 2 of their papers we have counts for

collaborators

7 papers

cs.CL2022

Fixing Model Bugs with Natural Language Patches

Shikhar Murty, Christopher D. Manning, Scott Lundberg +1

Current approaches for fixing systematic problems in NLP models (e.g. regex patches, finetuning on more data) are either brittle, or labor-intensive and liable to shortcuts. In con…

cs.CL20221 cited

ExSum: From Local Explanations to Model Understanding

Yilun Zhou, Marco Tulio Ribeiro, Julie Shah

Interpretability methods are developed to understand the working mechanisms of black-box models, which is crucial to their responsible deployment. Fulfilling this goal requires bot…

cs.CL2021

Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models

Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer +1

While counterfactual examples are useful for analysis and training of NLP models, current generation methods either rely on manual labor to create very few counterfactuals, or only…

cs.AI2020

Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance

Gagan Bansal, Tongshuang Wu, Joyce Zhou +5

Many researchers motivate explainable AI with studies showing that human-AI team performance on decision-making tasks improves when the AI explains its recommendations. However, pr…

cs.CL2020

Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin +1

Although measuring held-out accuracy has been the primary approach to evaluate generalization, it often overestimates the performance of NLP models, while alternative approaches fo…

cs.CV2020

SQuINTing at VQA Models: Introspecting VQA Models with Sub-Questions

Ramprasaath R. Selvaraju, Purva Tendulkar, Devi Parikh +4

Existing VQA datasets contain questions with varying levels of complexity. While the majority of questions in these datasets require perception for recognizing existence, propertie…