Offline bilingual word vectors, orthogonal transformations and the inverted softmax
arXiv:1702.03859
Abstract
Usually bilingual word vectors are trained "online". Mikolov et al. showed they can also be found "offline", whereby two pre-trained embeddings are aligned with a linear transformation, using dictionaries compiled from expert knowledge. In this work, we prove that the linear transformation between two spaces should be orthogonal. This transformation can be obtained using the singular value decomposition. We introduce a novel "inverted softmax" for identifying translation pairs, with which we improve the precision @1 of Mikolov's original mapping from 34% to 43%, when translating a test set composed of both common and rare English words into Italian. Orthogonal transformations are more robust to noise, enabling us to learn the transformation without expert bilingual signal by constructing a "pseudo-dictionary" from the identical character strings which appear in both languages, achieving 40% precision on the same test set. Finally, we extend our method to retrieve the true translations of English sentences from a corpus of 200k Italian sentences with a precision @1 of 68%.
Accepted to conference track at ICLR 2017
Cited by in corpus (24)
- Multilingual Alignment of Contextual Word Representations
- From Zero to Hero: On the Limitations of Zero-Shot Cross-Lingual Transfer with Multilingual Transformers
- Cross-Lingual Alignment of Contextual Word Embeddings, with Applications to Zero-shot Dependency Parsing
- Cross-lingual Alignment vs Joint Training: A Comparative Study and A Simple Unified Framework
- Bilingual Dictionary Based Neural Machine Translation without Using Parallel Sentences
- A Study of Cross-Lingual Ability and Language-specific Information in Multilingual BERT
- Two-Level Transformer and Auxiliary Coherence Modeling for Improved Text Segmentation
- Zero-Shot Cross-Lingual Opinion Target Extraction
- Do Explicit Alignments Robustly Improve Multilingual Encoders?
- Aligning Vector-spaces with Noisy Supervised Lexicons
- Cross-lingual Data Transformation and Combination for Text Classification
- HAL: Improved Text-Image Matching by Mitigating Visual Semantic Hubs
- Weakly-Supervised Concept-based Adversarial Learning for Cross-lingual Word Embeddings
- Expanding the Text Classification Toolbox with Cross-Lingual Embeddings
- Why Overfitting Isn't Always Bad: Retrofitting Cross-Lingual Word Embeddings to Dictionaries
- Sentence transition matrix: An efficient approach that preserves sentence semantics
- A Strong and Robust Baseline for Text-Image Matching
- Are We Consistently Biased? Multidimensional Analysis of Biases in Distributional Word Vectors
- Bi-Decoder Augmented Network for Neural Machine Translation
- Improving Unsupervised Word-by-Word Translation with Language Model and Denoising Autoencoder
- Document Network Embedding: Coping for Missing Content and Missing Links
- Revisiting Adversarial Autoencoder for Unsupervised Word Translation with Cycle Consistency and Improved Training
- Word Embeddings: Stability and Semantic Change
- RPD: A Distance Function Between Word Embeddings