A Survey Of Cross-lingual Word Embedding Models
arXiv:1706.04902 · doi:10.1613/jair.1.11640
Abstract
Cross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we provide a comprehensive typology of cross-lingual word embedding models. We compare their data requirements and objective functions. The recurring theme of the survey is that many of the models presented in the literature optimize for the same objectives, and that seemingly different models are often equivalent modulo optimization strategies, hyper-parameters, and such. We also discuss the different ways cross-lingual word embeddings are evaluated, as well as future challenges and research horizons.
Published in Journal of Artificial Intelligence Research
References in corpus (15)
- Distributed Representations of Sentences and Documents
- Introduction to the CoNLL-2002 Shared Task: Language-Independent Named Entity Recognition
- Recurrent Neural Network for Text Classification with Multi-Task Learning
- A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
- Unsupervised Machine Translation Using Monolingual Corpora Only
- A Network-based End-to-End Trainable Task-oriented Dialogue System
- Analyzing the Limitations of Cross-lingual Word Embedding Mappings
- Bilingual Lexicon Induction through Unsupervised Machine Translation
- Cross-lingual Models of Word Embeddings: An Empirical Comparison
- Semantic Specialisation of Distributional Word Vector Spaces using Monolingual and Cross-Lingual Constraints
- Word Embeddings and Their Use In Sentence Classification Tasks
- Learning Crosslingual Word Embeddings without Bilingual Corpora
- Multilingual Multi-modal Embeddings for Natural Language Processing
- Model Transfer for Tagging Low-resource Languages using a Bilingual Dictionary
- Sharing Network Parameters for Crosslingual Named Entity Recognition
Cited by in corpus (13)
- COVID-19 sentiment analysis via deep learning during the rise of novel cases
- Cultural Cartography with Word Embeddings
- Abstractive Text Summarization: State of the Art, Challenges, and Improvements
- Improving Classifier Training Efficiency for Automatic Cyberbullying Detection with Feature Density
- Corpus-Based Paraphrase Detection Experiments and Review
- Learning Backward Compatible Embeddings
- From Bytes to Biases: Investigating the Cultural Self-Perception of Large Language Models
- Compressing and Interpreting Word Embeddings with Latent Space Regularization and Interactive Semantics Probing
- Cross-Language Learning for Entity Matching
- Linear Transformations for Cross-lingual Sentiment Analysis
- The Impact of Cross-Lingual Adjustment of Contextual Word Representations on Zero-Shot Transfer
- Shaping the Future of Endangered and Low-Resource Languages -- Our Role in the Age of LLMs: A Keynote at ECIR 2024
- Searching Personal Collections