Generating Natural Adversarial Examples
arXiv:1710.11342
Abstract
Due to their complex nature, it is hard to characterize the ways in which machine learning models can misbehave or be exploited when deployed. Recent work on adversarial examples, i.e. inputs with minor perturbations that result in substantially different model predictions, is helpful in evaluating the robustness of these models by exposing the adversarial scenarios where they fail. However, these malicious perturbations are often unnatural, not semantically meaningful, and not applicable to complicated domains such as language. In this paper, we propose a framework to generate natural and legible adversarial examples that lie on the data manifold, by searching in semantic space of dense and continuous data representation, utilizing the recent advances in generative adversarial networks. We present generated adversaries to demonstrate the potential of the proposed approach for black-box classifiers for a wide range of applications such as image classification, textual entailment, and machine translation. We include experiments to show that the generated adversaries are natural, legible to humans, and useful in evaluating and analyzing black-box classifiers.
Published as a conference paper at the International Conference on Learning Representations (ICLR) 2018
References in corpus (5)
Cited by in corpus (119)
- BAE: BERT-based Adversarial Examples for Text Classification
- Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
- Adversarial Examples: Attacks and Defenses for Deep Learning
- A General Framework for Adversarial Examples with Objectives
- On Adversarial Examples for Character-Level Neural Machine Translation
- AdvPC: Transferable Adversarial Perturbations on 3D Point Clouds
- Using Attribution to Decode Dataset Bias in Neural Network Models for Chemistry
- Constructing Unrestricted Adversarial Examples with Generative Models
- Generative Data Augmentation for Commonsense Reasoning
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
- Reevaluating Adversarial Examples in Natural Language
- On the Generalizability of Neural Program Models with respect to Semantic-Preserving Program Transformations
- Adversarial Training against Location-Optimized Adversarial Patches
- Adv-BERT: BERT is not robust on misspellings! Generating nature adversarial samples on BERT
- HotFlip: White-Box Adversarial Examples for Text Classification
- OpenAttack: An Open-source Textual Adversarial Attack Toolkit
- Seq2Sick: Evaluating the Robustness of Sequence-to-Sequence Models with Adversarial Examples
- Controllable Unsupervised Text Attribute Transfer via Editing Entangled Latent Representation
- Shilling Black-box Recommender Systems by Learning to Generate Fake User Profiles
- Security and Privacy Issues in Deep Learning
- Attack Graph Convolutional Networks by Adding Fake Nodes
- Robust in Practice: Adversarial Attacks on Quantum Machine Learning
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- SemMT: A Semantic-based Testing Approach for Machine Translation Systems
- Defense Methods Against Adversarial Examples for Recurrent Neural Networks
- Fooling OCR Systems with Adversarial Text Images
- Robustness Verification for Transformers
- Adversarial Examples Make Strong Poisons
- On the Art and Science of Machine Learning Explanations
- How Should Pre-Trained Language Models Be Fine-Tuned Towards Adversarial Robustness?
- Security Matters: A Survey on Adversarial Machine Learning
- When and How to Fool Explainable Models (and Humans) with Adversarial Examples
- Towards a Robust Deep Neural Network in Texts: A Survey
- Improved Consistency Regularization for GANs
- Exposing Previously Undetectable Faults in Deep Neural Networks
- Combating Adversarial Misspellings with Robust Word Recognition
- DeepHunter: Hunting Deep Neural Network Defects via Coverage-Guided Fuzzing
- Generative Adversarial Networks: A Survey Towards Private and Secure Applications
- A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data Poisoning
- Explaining Deep Neural Networks
- Natural Backdoor Attack on Text Data
- An Empirical Survey of Data Augmentation for Limited Data Learning in NLP
- Towards Visual Distortion in Black-Box Attacks
- Generating 3D Adversarial Point Clouds
- Contextualized Perturbation for Textual Adversarial Attack
- Discrete Adversarial Attacks and Submodular Optimization with Applications to Text Classification
- Top-k Training of GANs: Improving GAN Performance by Throwing Away Bad Samples
- Metamorphic Relation Based Adversarial Attacks on Differentiable Neural Computer
- Adversarial Attacks and Defense on Texts: A Survey
- Chat as Expected: Learning to Manipulate Black-box Neural Dialogue Models
- PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification
- Adversarial Example Games
- Characterizing machine learning process: A maturity framework
- SoK: Machine Learning Governance
- A Reinforced Generation of Adversarial Examples for Neural Machine Translation
- Interpretable Adversarial Training for Text
- Defence against adversarial attacks using classical and quantum-enhanced Boltzmann machines
- Universal Adversarial Attacks with Natural Triggers for Text Classification
- Explainable Black-Box Attacks Against Model-based Authentication
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Generating Semantic Adversarial Examples via Feature Manipulation
- Evaluating Robustness to Context-Sensitive Feature Perturbations of Different Granularities
- ASR-GLUE: A New Multi-task Benchmark for ASR-Robust Natural Language Understanding
- A Survey on Out-of-Distribution Evaluation of Neural NLP Models
- Evaluation of Generalizability of Neural Program Analyzers under Semantic-Preserving Transformations
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- Practical Fast Gradient Sign Attack against Mammographic Image Classifier
- Universal, transferable and targeted adversarial attacks
- DQI: Measuring Data Quality in NLP
- GAP++: Learning to generate target-conditioned adversarial examples
- Detecting Operational Adversarial Examples for Reliable Deep Learning
- Identify Susceptible Locations in Medical Records via Adversarial Attacks on Deep Predictive Models
- CAGFuzz: Coverage-Guided Adversarial Generative Fuzzing Testing of Deep Learning Systems
- Semantics Preserving Adversarial Learning
- Exploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis
- Efficient Combinatorial Optimization for Word-level Adversarial Textual Attack
- Generalizable Adversarial Attacks with Latent Variable Perturbation Modelling
- Poison Attacks against Text Datasets with Conditional Adversarially Regularized Autoencoder
- Generating Out of Distribution Adversarial Attack using Latent Space Poisoning
- Are you tough enough? Framework for Robustness Validation of Machine Comprehension Systems
- Improving Robustness and Generality of NLP Models Using Disentangled Representations
- Effective and Imperceptible Adversarial Textual Attack via Multi-objectivization
- Harmonic Adversarial Attack Method
- Preserving Semantics in Textual Adversarial Attacks
- Structure-Preserving Transformation: Generating Diverse and Transferable Adversarial Examples
- T3: Tree-Autoencoder Constrained Adversarial Text Generation for Targeted Attack
- Model-Based Robust Deep Learning: Generalizing to Natural, Out-of-Distribution Data
- Customizing an Adversarial Example Generator with Class-Conditional GANs
- Internal Wasserstein Distance for Adversarial Attack and Defense
- Adversarial Attack and Defense of Structured Prediction Models
- Auditing AI models for Verified Deployment under Semantic Specifications
- Analysis Methods in Neural Language Processing: A Survey
- Generating Semantically Valid Adversarial Questions for TableQA
- AdvFoolGen: Creating Persistent Troubles for Deep Classifiers
- A cost-effective method for improving and re-purposing large, pre-trained GANs by fine-tuning their class-embeddings
- Lost In Translation: Generating Adversarial Examples Robust to Round-Trip Translation
- An Adversarially-Learned Turing Test for Dialog Generation Models
- Towards Variable-Length Textual Adversarial Attacks
- Distributional Robustness with IPMs and links to Regularization and GANs
- Data Augmentation via Structured Adversarial Perturbations
- Visual Attack and Defense on Text
- Adversarial Learning with Cost-Sensitive Classes
- The Vulnerability of the Neural Networks Against Adversarial Examples in Deep Learning Algorithms
- Towards Imperceptible Query-limited Adversarial Attacks with Perceptual Feature Fidelity Loss
- Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts
- Learning Multi-level Dependencies for Robust Word Recognition
- Challenge AI Mind: A Crowd System for Proactive AI Testing
- Adversarial Attacks Against Deep Learning Systems for ICD-9 Code Assignment
- Token Drop mechanism for Neural Machine Translation
- Perturbing Inputs for Fragile Interpretations in Deep Natural Language Processing
- Teaching Syntax by Adversarial Distraction
- Attack to Fool and Explain Deep Networks
- On the Transferability of Adversarial Attacksagainst Neural Text Classifier
- Games for Fairness and Interpretability
- Target Model Agnostic Adversarial Attacks with Query Budgets on Language Understanding Models
- Synthesizing Unrestricted False Positive Adversarial Objects Using Generative Models
- Human Imperceptible Attacks and Applications to Improve Fairness
- Robustness to Programmable String Transformations via Augmented Abstract Training
- Natural Adversarial Sentence Generation with Gradient-based Perturbation