"That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial Attacks
arXiv:2204.04636 · doi:10.18653/v1/2022.acl-long.538
Abstract
Adversarial attacks are a major challenge faced by current machine learning research. These purposely crafted inputs fool even the most advanced models, precluding their deployment in safety-critical applications. Extensive research in computer vision has been carried to develop reliable defense strategies. However, the same issue remains less explored in natural language processing. Our work presents a model-agnostic detector of adversarial text examples. The approach identifies patterns in the logits of the target classifier when perturbing the input text. The proposed detector improves the current state-of-the-art performance in recognizing adversarial inputs and exhibits strong generalization capabilities across different NLP models, datasets, and word-level attacks.
ACL 2022
References in corpus (6)
- Explaining and Harnessing Adversarial Examples
- Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
- Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
- Early Methods for Detecting Adversarial Images
- Towards Robustness Against Natural Language Word Substitutions
- Defense against Adversarial Attacks in NLP via Dirichlet Neighborhood Ensemble