Preserving Semantics in Textual Adversarial Attacks
arXiv:2211.04205 · doi:10.3233/FAIA230376
Abstract
The growth of hateful online content, or hate speech, has been associated with a global increase in violent crimes against minorities [23]. Harmful online content can be produced easily, automatically and anonymously. Even though, some form of auto-detection is already achieved through text classifiers in NLP, they can be fooled by adversarial attacks. To strengthen existing systems and stay ahead of attackers, we need better adversarial attacks. In this paper, we show that up to 70% of adversarial examples generated by adversarial attacks should be discarded because they do not preserve semantics. We address this core weakness and propose a new, fully supervised sentence embedding technique called Semantics-Preserving-Encoder (SPE). Our method outperforms existing sentence encoders used in adversarial attacks by achieving 1.2x - 5.1x better real attack success rate. We release our code as a plugin that can be used in any existing adversarial attack to improve its quality and speed up its execution.
8 pages, 4 figures
References in corpus (11)
- Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
- BAE: BERT-based Adversarial Examples for Text Classification
- Adversarial Machine Learning at Scale
- TextBugger: Generating Adversarial Text Against Real-world Applications
- Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
- Reevaluating Adversarial Examples in Natural Language
- Towards Robustness Against Natural Language Word Substitutions
- Contextualized Perturbation for Textual Adversarial Attack
- Robust Encodings: A Framework for Combating Adversarial Typos
- Model Extraction and Adversarial Transferability, Your BERT is Vulnerable!
- T3: Tree-Autoencoder Constrained Adversarial Text Generation for Targeted Attack