RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking
arXiv:2110.07367
Abstract
In various natural language processing tasks, passage retrieval and passage re-ranking are two key procedures in finding and ranking relevant information. Since both the two procedures contribute to the final performance, it is important to jointly optimize them in order to achieve mutual improvement. In this paper, we propose a novel joint training approach for dense passage retrieval and passage re-ranking. A major contribution is that we introduce the dynamic listwise distillation, where we design a unified listwise training approach for both the retriever and the re-ranker. During the dynamic distillation, the retriever and the re-ranker can be adaptively improved according to each other's relevance information. We also propose a hybrid data augmentation strategy to construct diverse training instances for listwise training approach. Extensive experiments show the effectiveness of our approach on both MSMARCO and Natural Questions datasets. Our code is available at https://github.com/PaddlePaddle/RocketQA.
EMNLP 2021
References in corpus (11)
- REALM: Retrieval-Augmented Language Model Pre-Training
- Deeper Text Understanding for IR with Contextual Neural Language Modeling
- Multi-Stage Document Ranking with BERT
- An Information Retrieval Approach to Short Text Conversation
- Understanding the Behaviors of BERT in Ranking
- Pre-training Tasks for Embedding-based Large-scale Retrieval
- RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering
- RepBERT: Contextualized Text Embeddings for First-Stage Retrieval
- Generation-Augmented Retrieval for Open-domain Question Answering
- Neural Passage Retrieval with Improved Negative Contrast
- Neural Retrieval for Question Answering with Cross-Attention Supervised Data Augmentation