Reevaluating Adversarial Examples in Natural Language
arXiv:2004.14174 · doi:10.18653/v1/2020.findings-emnlp.341
Abstract
State-of-the-art attacks on NLP models lack a shared definition of a what constitutes a successful attack. We distill ideas from past work into a unified framework: a successful natural language adversarial example is a perturbation that fools the model and follows some linguistic constraints. We then analyze the outputs of two state-of-the-art synonym substitution attacks. We find that their perturbations often do not preserve semantics, and 38% introduce grammatical errors. Human surveys reveal that to successfully preserve semantics, we need to significantly increase the minimum cosine similarities between the embeddings of swapped words and between the sentence encodings of original and perturbed sentences.With constraints adjusted to better preserve semantics and grammaticality, the attack success rate drops by over 70 percentage points.
15 pages; 9 Tables; 5 Figures
References in corpus (4)
Cited by in corpus (10)
- Testing the limits of natural language models for predicting human language judgments
- Contextualized Perturbation for Textual Adversarial Attack
- Searching for a Search Method: Benchmarking Search Algorithms for Generating NLP Adversarial Examples
- Generating Adversarial Examples in Chinese Texts Using Sentence-Pieces
- Towards Improving Adversarial Training of NLP Models
- Preserving Semantics in Textual Adversarial Attacks
- Grey-box Adversarial Attack And Defence For Sentiment Classification
- Gradient-based Adversarial Attacks against Text Transformers
- Towards Variable-Length Textual Adversarial Attacks
- Searching for an Effective Defender: Benchmarking Defense against Adversarial Word Substitution