Bilingual Lexicon Induction through Unsupervised Machine Translation
arXiv:1907.10761 · doi:10.18653/v1/P19-1494
Abstract
A recent research line has obtained strong results on bilingual lexicon induction by aligning independently trained word embeddings in two languages and using the resulting cross-lingual embeddings to induce word translation pairs through nearest neighbor or related retrieval methods. In this paper, we propose an alternative approach to this problem that builds on the recent work on unsupervised machine translation. This way, instead of directly inducing a bilingual lexicon from cross-lingual embeddings, we use them to build a phrase-table, combine it with a language model, and use the resulting machine translation system to generate a synthetic parallel corpus, from which we extract the bilingual lexicon using statistical word alignment techniques. As such, our method can work with any word embedding and cross-lingual mapping technique, and it does not require any additional resource besides the monolingual corpus used to train the embeddings. When evaluated on the exact same cross-lingual embeddings, our proposed method obtains an average improvement of 6 accuracy points over nearest neighbor and 4 points over CSLS retrieval, establishing a new state-of-the-art in the standard MUSE dataset.
ACL 2019
Cited by in corpus (13)
- On the Cross-lingual Transferability of Monolingual Representations
- A Call for More Rigor in Unsupervised Cross-lingual Learning
- Cross-lingual Retrieval for Iterative Self-Supervised Training
- A Survey of Orthographic Information in Machine Translation
- Cross-lingual Alignment vs Joint Training: A Comparative Study and A Simple Unified Framework
- Combining Static Word Embeddings and Contextual Representations for Bilingual Lexicon Induction
- On a Novel Application of Wasserstein-Procrustes for Unsupervised Cross-Lingual Learning
- Beyond Offline Mapping: Learning Cross Lingual Word Embeddings through Context Anchoring
- Bilingual Lexicon Induction via Unsupervised Bitext Construction and Word Alignment
- Predicting Performance for Natural Language Processing Tasks
- From meaning to perception -- exploring the space between word and odor perception embeddings
- A Relaxed Matching Procedure for Unsupervised BLI
- Data Augmentation with Unsupervised Machine Translation Improves the Structural Similarity of Cross-lingual Word Embeddings