Towards Multi-Sense Cross-Lingual Alignment of Contextual Embeddings
arXiv:2103.06459
Abstract
Cross-lingual word embeddings (CLWE) have been proven useful in many cross-lingual tasks. However, most existing approaches to learn CLWE including the ones with contextual embeddings are sense agnostic. In this work, we propose a novel framework to align contextual embeddings at the sense level by leveraging cross-lingual signal from bilingual dictionaries only. We operationalize our framework by first proposing a novel sense-aware cross entropy loss to model word senses explicitly. The monolingual ELMo and BERT models pretrained with our sense-aware cross entropy loss demonstrate significant performance improvement for word sense disambiguation tasks. We then propose a sense alignment objective on top of the sense-aware cross entropy loss for cross-lingual model pretraining, and pretrain cross-lingual models for several language pairs (English to German/Spanish/Japanese/Chinese). Compared with the best baseline results, our cross-lingual models achieve 0.52%, 2.09% and 1.29% average performance improvements on zero-shot cross-lingual NER, sentiment classification and XNLI tasks, respectively.
Accepted by COLING 2022
References in corpus (9)
- Multilingual Denoising Pre-training for Neural Machine Translation
- Word Translation Without Parallel Data
- Multilingual Alignment of Contextual Word Representations
- How multilingual is Multilingual BERT?
- Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT
- Data Stream Clustering: Challenges and Issues
- Cross-Lingual Alignment of Contextual Word Embeddings, with Applications to Zero-shot Dependency Parsing
- Cross-Lingual Contextual Word Embeddings Mapping With Multi-Sense Words In Mind
- Adversarial Learning with Contextual Embeddings for Zero-resource Cross-lingual Classification and NER