Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
arXiv:1901.06796
Abstract
With the development of high computational devices, deep neural networks (DNNs), in recent years, have gained significant popularity in many Artificial Intelligence (AI) applications. However, previous efforts have shown that DNNs were vulnerable to strategically modified samples, named adversarial examples. These samples are generated with some imperceptible perturbations but can fool the DNNs to give false predictions. Inspired by the popularity of generating adversarial examples for image DNNs, research efforts on attacking DNNs for textual applications emerges in recent years. However, existing perturbation methods for images cannotbe directly applied to texts as text data is discrete. In this article, we review research works that address this difference and generatetextual adversarial examples on DNNs. We collect, select, summarize, discuss and analyze these works in a comprehensive way andcover all the related information to make the article self-contained. Finally, drawing on the reviewed literature, we provide further discussions and suggestions on this topic.
40
References in corpus (10)
- Distilling the Knowledge in a Neural Network
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Convolutional Neural Networks for Sentence Classification
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
- Towards Crafting Text Adversarial Samples
- Black-Box Attacks against RNN based Malware Detection Algorithms
- Adversarial Texts with Gradient Methods
- DANCin SEQ2SEQ: Fooling Text Classifiers with Adversarial Text Example Generation
Cited by in corpus (10)
- Deep Interest Highlight Network for Click-Through Rate Prediction in Trigger-Induced Recommendation
- Beyond Leaderboards: A survey of methods for revealing weaknesses in Natural Language Inference data and models
- Chat as Expected: Learning to Manipulate Black-box Neural Dialogue Models
- On Adversarial Examples for Biomedical NLP Tasks
- Stress Test Evaluation of Biomedical Word Embeddings
- Risk Management Framework for Machine Learning Security
- Deceptive Deletions for Protecting Withdrawn Posts on Social Platforms
- Contrastive Fine-tuning Improves Robustness for Neural Rankers
- Gödel's Sentence Is An Adversarial Example But Unsolvable
- Adversarial Attacks on Deep Models for Financial Transaction Records