Seq2Sick: Evaluating the Robustness of Sequence-to-Sequence Models with Adversarial Examples
arXiv:1803.01128
Abstract
Crafting adversarial examples has become an important technique to evaluate the robustness of deep neural networks (DNNs). However, most existing works focus on attacking the image classification problem since its input space is continuous and output space is finite. In this paper, we study the much more challenging problem of crafting adversarial examples for sequence-to-sequence (seq2seq) models, whose inputs are discrete text strings and outputs have an almost infinite number of possibilities. To address the challenges caused by the discrete input space, we propose a projected gradient method combined with group lasso and gradient regularization. To handle the almost infinite output space, we design some novel loss functions to conduct non-overlapping attack and targeted keyword attack. We apply our algorithm to machine translation and text summarization tasks, and verify the effectiveness of the proposed algorithm: by changing less than 3 words, we can make seq2seq model to produce desired outputs with high success rates. On the other hand, we recognize that, compared with the well-evaluated CNN-based classifiers, seq2seq models are intrinsically more robust to adversarial attacks.
References in corpus (10)
- Explaining and Harnessing Adversarial Examples
- Understanding Neural Networks through Representation Erasure
- Generating Natural Adversarial Examples
- Toward Multilingual Neural Machine Translation with Universal Encoder and Decoder
- Adversarial Examples for Evaluating Reading Comprehension Systems
- Towards Crafting Text Adversarial Samples
- HotFlip: White-Box Adversarial Examples for Text Classification
- Adversarial Texts with Gradient Methods
- Greedy Attack and Gumbel Attack: Generating Adversarial Examples for Discrete Data
- On Evaluation of Adversarial Perturbations for Sequence-to-Sequence Models
Cited by in corpus (23)
- TextBugger: Generating Adversarial Text Against Real-world Applications
- Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
- Query-Efficient Hard-label Black-box Attack:An Optimization-based Approach
- Reevaluating Adversarial Examples in Natural Language
- TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP
- Attack Graph Convolutional Networks by Adding Fake Nodes
- Defense Methods Against Adversarial Examples for Recurrent Neural Networks
- Defend Deep Neural Networks Against Adversarial Examples via Fixed and Dynamic Quantized Activation Functions
- Explaining Deep Neural Networks
- Exploiting Rich Syntactic Information for Semantic Parsing with Graph-to-Sequence Model
- Say What I Want: Towards the Dark Side of Neural Dialogue Models
- Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation
- On Evaluation of Adversarial Perturbations for Sequence-to-Sequence Models
- Chat as Expected: Learning to Manipulate Black-box Neural Dialogue Models
- Generating Natural Language Adversarial Examples on a Large Scale with Generative Models
- A Reinforced Generation of Adversarial Examples for Neural Machine Translation
- Query-Efficient Black-Box Attack by Active Learning
- T3: Tree-Autoencoder Constrained Adversarial Text Generation for Targeted Attack
- Analysis Methods in Neural Language Processing: A Survey
- Is Ordered Weighted Regularized Regression Robust to Adversarial Perturbation? A Case Study on OSCAR
- Challenge AI Mind: A Crowd System for Proactive AI Testing
- Knowing When to Stop: Evaluation and Verification of Conformity to Output-size Specifications
- Natural Adversarial Sentence Generation with Gradient-based Perturbation