BAE: BERT-based Adversarial Examples for Text Classification
arXiv:2004.01970 · doi:10.18653/v1/2020.emnlp-main.498
Abstract
Modern text classification models are susceptible to adversarial examples, perturbed versions of the original text indiscernible by humans which get misclassified by the model. Recent works in NLP use rule-based synonym replacement strategies to generate adversarial examples. These strategies can lead to out-of-context and unnaturally complex token replacements, which are easily identifiable by humans. We present BAE, a black box attack for generating adversarial examples using contextual perturbations from a BERT masked language model. BAE replaces and inserts tokens in the original text by masking a portion of the text and leveraging the BERT-MLM to generate alternatives for the masked tokens. Through automatic and human evaluations, we show that BAE performs a stronger attack, in addition to generating adversarial examples with improved grammaticality and semantic coherence as compared to prior work.
Accepted at EMNLP 2020 Main Conference
References in corpus (2)
Cited by in corpus (26)
- On the Opportunities and Risks of Foundation Models
- Secure and Trustworthy Artificial Intelligence-Extended Reality (AI-XR) for Metaverses
- Semantic Robustness of Models of Source Code
- OpenAttack: An Open-source Textual Adversarial Attack Toolkit
- Can Adversarial Weight Perturbations Inject Neural Backdoors?
- TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP
- Contextualized Perturbation for Textual Adversarial Attack
- Better Robustness by More Coverage: Adversarial Training with Mixup Augmentation for Robust Fine-tuning
- SoK: Machine Learning Governance
- Defense against adversarial attacks on deep convolutional neural networks through nonlocal denoising
- Generating Adversarial Examples in Chinese Texts Using Sentence-Pieces
- Adversarial Training with Contrastive Learning in NLP
- Towards Improving Adversarial Training of NLP Models
- A Closer Look into Transformer-Based Code Intelligence Through Code Transformation: Challenges and Opportunities
- Preserving Semantics in Textual Adversarial Attacks
- How Vulnerable Are Automatic Fake News Detection Methods to Adversarial Attacks?
- Code-Mixing on Sesame Street: Dawn of the Adversarial Polyglots
- Gradient-based Adversarial Attacks against Text Transformers
- R&R: Metric-guided Adversarial Sentence Generation
- Adversarial Evaluation of Multimodal Models under Realistic Gray Box Assumption
- A Sweet Rabbit Hole by DARCY: Using Honeypots to Detect Universal Trigger's Adversarial Attacks
- Improved and Efficient Text Adversarial Attacks using Target Information
- Enhancing Interpretable Clauses Semantically using Pretrained Word Representation
- Improving Zero-Shot Cross-Lingual Transfer Learning via Robust Training
- Text Counterfactuals via Latent Optimization and Shapley-Guided Search
- On the Transferability of Adversarial Attacksagainst Neural Text Classifier