It's Morphin' Time! Combating Linguistic Discrimination with Inflectional Perturbations
arXiv:2005.04364 · doi:10.18653/v1/2020.acl-main.263
Abstract
Training on only perfect Standard English corpora predisposes pre-trained neural networks to discriminate against minorities from non-standard linguistic backgrounds (e.g., African American Vernacular English, Colloquial Singapore English, etc.). We perturb the inflectional morphology of words to craft plausible and semantically similar adversarial examples that expose these biases in popular NLP models, e.g., BERT and Transformer, and show that adversarially fine-tuning them for a single epoch significantly improves robustness without sacrificing performance on clean data.
To appear in the Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020)
References in corpus (3)
Cited by in corpus (19)
- Post-hoc Interpretability for Neural NLP: A Survey
- TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP
- Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures
- Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models
- An Empirical Survey of Data Augmentation for Limited Data Learning in NLP
- Better Robustness by More Coverage: Adversarial Training with Mixup Augmentation for Robust Fine-tuning
- Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
- Effective and Imperceptible Adversarial Textual Attack via Multi-objectivization
- Efficient Combinatorial Optimization for Word-level Adversarial Textual Attack
- From Hero to Zéroe: A Benchmark of Low-Level Adversarial Attacks
- Generating Syntactically Controlled Paraphrases without Using Annotated Parallel Pairs
- Addressing the Vulnerability of NMT in Input Perturbations
- Code-Mixing on Sesame Street: Dawn of the Adversarial Polyglots
- CLINE: Contrastive Learning with Semantic Negative Examples for Natural Language Understanding
- Spinning Sequence-to-Sequence Models with Meta-Backdoors
- Achieving Model Robustness through Discrete Adversarial Training
- On the Robustness of Intent Classification and Slot Labeling in Goal-oriented Dialog Systems to Real-world Noise
- Evaluating the Morphosyntactic Well-formedness of Generated Texts
- On the Universality of Deep Contextual Language Models