Cross-Lingual Training with Dense Retrieval for Document Retrieval
arXiv:2109.01628
Abstract
Dense retrieval has shown great success in passage ranking in English. However, its effectiveness in document retrieval for non-English languages remains unexplored due to the limitation in training resources. In this work, we explore different transfer techniques for document ranking from English annotations to multiple non-English languages. Our experiments on the test collections in six languages (Chinese, Arabic, French, Hindi, Bengali, Spanish) from diverse language families reveal that zero-shot model-based transfer using mBERT improves the search quality in non-English mono-lingual retrieval. Also, we find that weakly-supervised target language transfer yields competitive performances against the generation-based target language transfer that requires external translators and query generators.
References in corpus (7)
- Multilingual Denoising Pre-training for Neural Machine Translation
- Deeper Text Understanding for IR with Contextual Neural Language Modeling
- Simple Applications of BERT for Ad Hoc Document Retrieval
- Pre-training Tasks for Embedding-based Large-scale Retrieval
- A Study of Neural Matching Models for Cross-lingual IR
- Cross-Lingual Relevance Transfer for Document Retrieval
- Cross-language Sentence Selection via Data Augmentation and Rationale Training