220 citations · 878 across the 42 of their papers we have counts for
Showing 2021 · cs.CLShow all
2 papers · 2 filters
cs.CL2021
Perturbing Inputs for Fragile Interpretations in Deep Natural Language Processing
Sanchit Sinha, Hanjie Chen, Arshdeep Sekhon +2
Interpretability methods like Integrated Gradient and LIME are popular choices for explaining natural language model predictions with relative word importance scores. These interpr…
cs.CL2021★ 5 cited
Towards Improving Adversarial Training of NLP Models
Jin Yong Yoo, Yanjun Qi
Adversarial training, a method for learning robust deep neural networks, constructs adversarial examples during training. However, recent methods for generating NLP adversarial exa…