Word Translation Without Parallel Data
arXiv:1710.04087
Abstract
State-of-the-art methods for learning cross-lingual word embeddings have relied on bilingual dictionaries or parallel corpora. Recent studies showed that the need for parallel data supervision can be alleviated with character-level information. While these methods showed encouraging results, they are not on par with their supervised counterparts and are limited to pairs of languages sharing a common alphabet. In this work, we show that we can build a bilingual dictionary between two languages without using any parallel corpora, by aligning monolingual word embedding spaces in an unsupervised way. Without using any character information, our model even outperforms existing supervised methods on cross-lingual tasks for some language pairs. Our experiments demonstrate that our method works very well also for distant language pairs, like English-Russian or English-Chinese. We finally describe experiments on the English-Esperanto low-resource language pair, on which there only exists a limited amount of parallel data, to show the potential impact of our method in fully unsupervised machine translation. Our code, embeddings and dictionaries are publicly available.
ICLR 2018
References in corpus (1)
Cited by in corpus (54)
- Multilingual Denoising Pre-training for Neural Machine Translation
- Does Object Recognition Work for Everyone?
- Multi-modal Sarcasm Detection and Humor Classification in Code-mixed Conversations
- XLDA: Cross-Lingual Data Augmentation for Natural Language Inference and Question Answering
- Cross-lingual Zero- and Few-shot Hate Speech Detection Utilising Frozen Transformer Language Models and AXEL
- A Study of Neural Matching Models for Cross-lingual IR
- Integrating Social Media into a Pan-European Flood Awareness System: A Multilingual Approach
- Non-Autoregressive Neural Machine Translation with Enhanced Decoder Input
- Cross-Lingual Named Entity Recognition Using Parallel Corpus: A New Approach Using XLM-RoBERTa Alignment
- Crosslingual Document Embedding as Reduced-Rank Ridge Regression
- QUEACO: Borrowing Treasures from Weakly-labeled Behavior Data for Query Attribute Value Extraction
- Deep Matching Autoencoders
- word2word: A Collection of Bilingual Lexicons for 3,564 Language Pairs
- Combining Pre-trained Word Embeddings and Linguistic Features for Sequential Metaphor Identification
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- PidginUNMT: Unsupervised Neural Machine Translation from West African Pidgin to English
- Updating Pre-trained Word Vectors and Text Classifiers using Monolingual Alignment
- Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer
- Cross-Lingual Transfer Learning for Question Answering
- A Robust Self-Learning Method for Fully Unsupervised Cross-Lingual Mappings of Word Embeddings: Making the Method Robustly Reproducible as Well
- Towards Zero-shot Cross-lingual Image Retrieval and Tagging
- Contextual Lensing of Universal Sentence Representations
- Compass-aligned Distributional Embeddings for Studying Semantic Differences across Corpora
- Translating the Unseen? Yoruba-English MT in Low-Resource, Morphologically-Unmarked Settings
- Aligning Vector-spaces with Noisy Supervised Lexicons
- Neural Decipherment via Minimum-Cost Flow: from Ugaritic to Linear B
- Revisiting Language Encoding in Learning Multilingual Representations
- A Survey on Low-Resource Neural Machine Translation
- Discovering Bilingual Lexicons in Polyglot Word Embeddings
- Towards Reducing Bias in Gender Classification
- Traceability Support for Multi-Lingual Software Projects
- Low-Resource Sequence Labeling via Unsupervised Multilingual Contextualized Representations
- The Typology of Polysemy: A Multilingual Distributional Framework
- Worse WER, but Better BLEU? Leveraging Word Embedding as Intermediate in Multitask End-to-End Speech Translation
- SAR: Learning Cross-Language API Mappings with Little Knowledge
- MaskParse@Deskin at SemEval-2019 Task 1: Cross-lingual UCCA Semantic Parsing using Recursive Masked Sequence Tagging
- DeepSubQE: Quality estimation for subtitle translations
- Word Embedding Transformation for Robust Unsupervised Bilingual Lexicon Induction
- Bilingual Dictionary-based Language Model Pretraining for Neural Machine Translation
- Unsupervised Transfer Learning in Multilingual Neural Machine Translation with Cross-Lingual Word Embeddings
- Sentence transition matrix: An efficient approach that preserves sentence semantics
- Multilingual, Temporal and Sentimental Distant-Reading of City Events
- Contrastive Language Adaptation for Cross-Lingual Stance Detection
- Event-Driven Query Expansion
- Unsupervised Word Translation Pairing using Refinement based Point Set Registration
- CLAR: A Cross-Lingual Argument Regularizer for Semantic Role Labeling
- Cross-lingual Word Embeddings beyond Zero-shot Machine Translation
- Cross-lingual transfer learning for spoken language understanding
- Language Model-Driven Unsupervised Neural Machine Translation
- Target-Oriented Fine-tuning for Zero-Resource Named Entity Recognition
- From Machine Translation to Code-Switching: Generating High-Quality Code-Switched Text
- Mapping Supervised Bilingual Word Embeddings from English to low-resource languages
- Learning aligned embeddings for semi-supervised word translation using Maximum Mean Discrepancy
- Wasserstein distances for evaluating cross-lingual embeddings