Offline bilingual word vectors, orthogonal transformations and the inverted softmax
arXiv:1702.03859
Abstract
Usually bilingual word vectors are trained "online". Mikolov et al. showed they can also be found "offline", whereby two pre-trained embeddings are aligned with a linear transformation, using dictionaries compiled from expert knowledge. In this work, we prove that the linear transformation between two spaces should be orthogonal. This transformation can be obtained using the singular value decomposition. We introduce a novel "inverted softmax" for identifying translation pairs, with which we improve the precision @1 of Mikolov's original mapping from 34% to 43%, when translating a test set composed of both common and rare English words into Italian. Orthogonal transformations are more robust to noise, enabling us to learn the transformation without expert bilingual signal by constructing a "pseudo-dictionary" from the identical character strings which appear in both languages, achieving 40% precision on the same test set. Finally, we extend our method to retrieve the true translations of English sentences from a corpus of 200k Italian sentences with a precision @1 of 68%.
Accepted to conference track at ICLR 2017
Cited by in corpus (29)
- Multilingual Alignment of Contextual Word Representations
- Time-aware Graph Neural Networks for Entity Alignment between Temporal Knowledge Graphs
- From Zero to Hero: On the Limitations of Zero-Shot Cross-Lingual Transfer with Multilingual Transformers
- Cross-Lingual Alignment of Contextual Word Embeddings, with Applications to Zero-shot Dependency Parsing
- Bilingual Dictionary Based Neural Machine Translation without Using Parallel Sentences
- A Study of Cross-Lingual Ability and Language-specific Information in Multilingual BERT
- Cross-lingual Alignment vs Joint Training: A Comparative Study and A Simple Unified Framework
- Two-Level Transformer and Auxiliary Coherence Modeling for Improved Text Segmentation
- Zero-Shot Cross-Lingual Opinion Target Extraction
- Do Explicit Alignments Robustly Improve Multilingual Encoders?
- Aligning Vector-spaces with Noisy Supervised Lexicons
- Cross-lingual Data Transformation and Combination for Text Classification
- HAL: Improved Text-Image Matching by Mitigating Visual Semantic Hubs
- Combining Static Word Embeddings and Contextual Representations for Bilingual Lexicon Induction
- Sentence transition matrix: An efficient approach that preserves sentence semantics
- Expanding the Text Classification Toolbox with Cross-Lingual Embeddings
- Why Overfitting Isn't Always Bad: Retrofitting Cross-Lingual Word Embeddings to Dictionaries
- Weakly-Supervised Concept-based Adversarial Learning for Cross-lingual Word Embeddings
- A Strong and Robust Baseline for Text-Image Matching
- Unsupervised Word Translation Pairing using Refinement based Point Set Registration
- Learning Efficient Task-Specific Meta-Embeddings with Word Prisms
- Examining Cross-lingual Contextual Embeddings with Orthogonal Structural Probes
- Bi-Decoder Augmented Network for Neural Machine Translation
- Document Network Embedding: Coping for Missing Content and Missing Links
- Improving Unsupervised Word-by-Word Translation with Language Model and Denoising Autoencoder
- Word Embeddings: Stability and Semantic Change
- Revisiting Adversarial Autoencoder for Unsupervised Word Translation with Cycle Consistency and Improved Training
- RPD: A Distance Function Between Word Embeddings
- Are We Consistently Biased? Multidimensional Analysis of Biases in Distributional Word Vectors