A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
arXiv:1805.06297 · doi:10.18653/v1/P18-1073
Abstract
Recent work has managed to learn cross-lingual word embeddings without parallel data by mapping monolingual embeddings to a shared space through adversarial training. However, their evaluation has focused on favorable conditions, using comparable corpora or closely-related languages, and we show that they often fail in more realistic scenarios. This work proposes an alternative approach based on a fully unsupervised initialization that explicitly exploits the structural similarity of the embeddings, and a robust self-learning algorithm that iteratively improves this solution. Our method succeeds in all tested scenarios and obtains the best published results in standard datasets, even surpassing previous supervised systems. Our implementation is released as an open source project at https://github.com/artetxem/vecmap
ACL 2018
References in corpus (1)
Cited by in corpus (52)
- On the Cross-lingual Transferability of Monolingual Representations
- Unsupervised Statistical Machine Translation
- An Effective Approach to Unsupervised Machine Translation
- Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
- Unsupervised Automatic Speech Recognition: A Review
- Unsupervised Hyperalignment for Multilingual Word Embeddings
- Analyzing the Limitations of Cross-lingual Word Embedding Mappings
- XLDA: Cross-Lingual Data Augmentation for Natural Language Inference and Question Answering
- Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT
- A Call for More Rigor in Unsupervised Cross-lingual Learning
- Bilingual Lexicon Induction through Unsupervised Machine Translation
- When Does Unsupervised Machine Translation Work?
- Unsupervised Neural Machine Translation Initialized by Unsupervised Statistical Machine Translation
- Zero-shot Reading Comprehension by Cross-lingual Transfer Learning with Multi-lingual Language Representation Model
- A Resource-Free Evaluation Metric for Cross-Lingual Word Embeddings Based on Graph Modularity
- Exploring Fine-tuning Techniques for Pre-trained Cross-lingual Models via Continual Learning
- On the Robustness of Unsupervised and Semi-supervised Cross-lingual Word Embedding Learning
- Robust Cross-lingual Embeddings from Parallel Sentences
- Revisiting the Context Window for Cross-lingual Word Embeddings
- Unsupervised Multilingual Alignment using Wasserstein Barycenter
- Learning Unsupervised Word Mapping by Maximizing Mean Discrepancy
- Filtered Inner Product Projection for Crosslingual Embedding Alignment
- Multi-Source Cross-Lingual Model Transfer: Learning What to Share
- Monolingual and Parallel Corpora for Kangri Low Resource Language
- Are Girls Neko or Shōjo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative Normalization
- Bilingual Lexicon Induction with Semi-supervision in Non-Isometric Embedding Spaces
- Towards Unsupervised Speech-to-Text Translation
- Learning Cross-lingual Embeddings from Twitter via Distant Supervision
- Modeling Named Entity Embedding Distribution into Hypersphere
- Unsupervised Translation of German--Lower Sorbian: Exploring Training and Novel Transfer Methods on a Low-Resource Language
- Unsupervised cross-lingual matching of product classifications
- Aligning Vector-spaces with Noisy Supervised Lexicons
- Unsupervised Hierarchy Matching with Optimal Transport over Hyperbolic Spaces
- Cross-lingual Entity Alignment with Incidental Supervision
- A Survey on Low-Resource Neural Machine Translation
- Interactive Refinement of Cross-Lingual Word Embeddings
- On a Novel Application of Wasserstein-Procrustes for Unsupervised Cross-Lingual Learning
- Density Matching for Bilingual Word Embedding
- Bilingual Language Modeling, A transfer learning technique for Roman Urdu
- Improving Candidate Generation for Low-resource Cross-lingual Entity Linking
- Multilingual Embeddings Jointly Induced from Contexts and Concepts: Simple, Strong and Scalable
- Duality Regularization for Unsupervised Bilingual Lexicon Induction
- Refinement of Unsupervised Cross-Lingual Word Embeddings
- Exploiting Cross-Lingual Subword Similarities in Low-Resource Document Classification
- Lost in Evaluation: Misleading Benchmarks for Bilingual Dictionary Induction
- Dynamical Pose Estimation
- Bot-Match: Social Bot Detection with Recursive Nearest Neighbors Search
- Explicit Cross-lingual Pre-training for Unsupervised Machine Translation
- Data Augmentation with Unsupervised Machine Translation Improves the Structural Similarity of Cross-lingual Word Embeddings
- Scrambled Translation Problem: A Problem of Denoising UNMT
- Open Named Entity Modeling from Embedding Distribution
- Unsupervised Clinical Language Translation