HotFlip: White-Box Adversarial Examples for Text Classification
arXiv:1712.06751
Abstract
We propose an efficient method to generate white-box adversarial examples to trick a character-level neural classifier. We find that only a few manipulations are needed to greatly decrease the accuracy. Our method relies on an atomic flip operation, which swaps one token for another, based on the gradients of the one-hot input vectors. Due to efficiency of our method, we can perform adversarial training which makes the model more robust to attacks at test time. With the use of a few semantics-preserving constraints, we demonstrate that HotFlip can be adapted to attack a word-level classifier as well.
References in corpus (4)
Cited by in corpus (44)
- Synthetic and Natural Noise Both Break Neural Machine Translation
- Pathologies of Neural Models Make Interpretations Difficult
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
- Reevaluating Adversarial Examples in Natural Language
- Adv-BERT: BERT is not robust on misspellings! Generating nature adversarial samples on BERT
- Seq2Sick: Evaluating the Robustness of Sequence-to-Sequence Models with Adversarial Examples
- Advbox: a toolbox to generate adversarial examples that fool neural networks
- L2-Nonexpansive Neural Networks
- MaskDGA: A Black-box Evasion Technique Against DGA Classifiers and Adversarial Defenses
- Self-Explaining Structures Improve NLP Models
- Natural Backdoor Attack on Text Data
- Discrete Adversarial Attacks and Submodular Optimization with Applications to Text Classification
- Certified Robustness to Adversarial Word Substitutions
- Learning to Discriminate Perturbations for Blocking Adversarial Attacks in Text Classification
- A Real-time Defense against Website Fingerprinting Attacks
- Efficient (Soft) Q-Learning for Text Generation with Limited Good Data
- Adversarial Attacks and Defense on Texts: A Survey
- Conditional BERT Contextual Augmentation
- SoK: Machine Learning Governance
- Universal Adversarial Perturbation for Text Classification
- Elephant in the Room: An Evaluation Framework for Assessing Adversarial Examples in NLP
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Defending Against Backdoor Attacks in Natural Language Generation
- Generating Adversarial Examples in Chinese Texts Using Sentence-Pieces
- Real-Time Adversarial Attacks
- Word Shape Matters: Robust Machine Translation with Visual Embedding
- Improving Robustness and Generality of NLP Models Using Disentangled Representations
- Generalizable Adversarial Attacks with Latent Variable Perturbation Modelling
- Random Directional Attack for Fooling Deep Neural Networks
- Poison Attacks against Text Datasets with Conditional Adversarially Regularized Autoencoder
- One Bit Matters: Understanding Adversarial Examples as the Abuse of Redundancy
- Position Bias Mitigation: A Knowledge-Aware Graph Model for Emotion Cause Extraction
- Grey-box Adversarial Attack And Defence For Sentiment Classification
- MALCOM: Generating Malicious Comments to Attack Neural Fake News Detection Models
- Systematic Attack Surface Reduction For Deployed Sentiment Analysis Models
- On Robustness and Bias Analysis of BERT-based Relation Extraction
- Towards Variable-Length Textual Adversarial Attacks
- An Adversarially-Learned Turing Test for Dialog Generation Models
- Towards Interpreting Recurrent Neural Networks through Probabilistic Abstraction
- Learning Multi-level Dependencies for Robust Word Recognition
- Leveraging Adversarial Training in Self-Learning for Cross-Lingual Text Classification
- Robustness Tests of NLP Machine Learning Models: Search and Semantically Replace
- Adversarially Robust and Explainable Model Compression with On-Device Personalization for Text Classification
- Adversarial Evaluation of Multimodal Models under Realistic Gray Box Assumption