Synthetic and Natural Noise Both Break Neural Machine Translation
arXiv:1711.02173
Abstract
Character-based neural machine translation (NMT) models alleviate out-of-vocabulary issues, learn morphology, and move us closer to completely end-to-end translation systems. Unfortunately, they are also very brittle and easily falter when presented with noisy data. In this paper, we confront NMT models with synthetic and natural sources of noise. We find that state-of-the-art models fail to translate even moderately noisy texts that humans have no trouble comprehending. We explore two approaches to increase model robustness: structure-invariant word representations and robust training on noisy texts. We find that a model based on a character convolutional neural network is able to simultaneously learn representations robust to multiple kinds of noise.
ICLR 2018 camera-ready
References in corpus (7)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Poisoning Attacks against Support Vector Machines
- Towards Crafting Text Adversarial Samples
- HotFlip: White-Box Adversarial Examples for Text Classification
- Black-box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers
- Large-Scale Machine Translation between Arabic and Hebrew: Available Corpora and Initial Results
Cited by in corpus (54)
- Shortcut Learning in Deep Neural Networks
- A Survey on Data Augmentation for Text Classification
- TextBugger: Generating Adversarial Text Against Real-world Applications
- Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
- Pathologies of Neural Models Make Interpretations Difficult
- Data Augmentation in Natural Language Processing: A Novel Text Generation Approach for Long and Short Text Classifiers
- Investigating Backtranslation in Neural Machine Translation
- WT5?! Training Text-to-Text Models to Explain their Predictions
- Generative Data Augmentation for Commonsense Reasoning
- It's Morphin' Time! Combating Linguistic Discrimination with Inflectional Perturbations
- CharBERT: Character-aware Pre-trained Language Model
- Adv-BERT: BERT is not robust on misspellings! Generating nature adversarial samples on BERT
- HotFlip: White-Box Adversarial Examples for Text Classification
- OpenAttack: An Open-source Textual Adversarial Attack Toolkit
- Adversarial NLI: A New Benchmark for Natural Language Understanding
- Adversarial Robustness of Deep Code Comment Generation
- Adversarially Regularising Neural NLI Models to Integrate Logical Background Knowledge
- Gradient Starvation: A Learning Proclivity in Neural Networks
- SemMT: A Semantic-based Testing Approach for Machine Translation Systems
- Improving Neural Machine Translation Robustness via Data Augmentation: Beyond Back Translation
- A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures
- When and How to Fool Explainable Models (and Humans) with Adversarial Examples
- A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data Poisoning
- Exploiting Rich Syntactic Information for Semantic Parsing with Graph-to-Sequence Model
- Discrete Adversarial Attacks and Submodular Optimization with Applications to Text Classification
- SGD-X: A Benchmark for Robust Generalization in Schema-Guided Dialogue Systems
- Certified Robustness to Adversarial Word Substitutions
- Improving the Robustness of Speech Translation
- Adversarial Attacks and Defense on Texts: A Survey
- A Reinforced Generation of Adversarial Examples for Neural Machine Translation
- Semantic Equivalent Adversarial Data Augmentation for Visual Question Answering
- Leveraging Automated Unit Tests for Unsupervised Code Translation
- Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation
- White-to-Black: Efficient Distillation of Black-Box Adversarial Attacks
- Large-scale Pretraining for Neural Machine Translation with Tens of Billions of Sentence Pairs
- Word Shape Matters: Robust Machine Translation with Visual Embedding
- BPE and CharCNNs for Translation of Morphology: A Cross-Lingual Comparison and Analysis
- PheMT: A Phenomenon-wise Dataset for Machine Translation Robustness on User-Generated Contents
- Contextual Text Denoising with Masked Language Models
- Visual Cues and Error Correction for Translation Robustness
- Stress Test Evaluation of Biomedical Word Embeddings
- Robust Neural Machine Translation with Joint Textual and Phonetic Embedding
- Can NMT Understand Me? Towards Perturbation-based Evaluation of NMT Models for Code Generation
- How to Select One Among All? An Extensive Empirical Study Towards the Robustness of Knowledge Distillation in Natural Language Understanding
- An unsupervised learning approach to evaluate questionnaire data -- what one can learn from violations of measurement invariance
- An Approach to Improve Robustness of NLP Systems against ASR Errors
- True or False: Does the Deep Learning Model Learn to Detect Rumors?
- Neural Machine Translation: A Review and Survey
- Robust Machine Translation with Domain Sensitive Pseudo-Sources: Baidu-OSU WMT19 MT Robustness Shared Task System Report
- Learning Multi-level Dependencies for Robust Word Recognition
- Generating Watermarked Adversarial Texts
- Adversarial Examples Generation for Reducing Implicit Gender Bias in Pre-trained Models
- Topic Model Robustness to Automatic Speech Recognition Errors in Podcast Transcripts
- Machine Translation in Pronunciation Space