B-PROP: Bootstrapped Pre-training with Representative Words Prediction for Ad-hoc Retrieval
arXiv:2104.09791
Abstract
Pre-training and fine-tuning have achieved remarkable success in many downstream natural language processing (NLP) tasks. Recently, pre-training methods tailored for information retrieval (IR) have also been explored, and the latest success is the PROP method which has reached new SOTA on a variety of ad-hoc retrieval benchmarks. The basic idea of PROP is to construct the \textit{representative words prediction} (ROP) task for pre-training inspired by the query likelihood model. Despite its exciting performance, the effectiveness of PROP might be bounded by the classical unigram language model adopted in the ROP task construction process. To tackle this problem, we propose a bootstrapped pre-training method (namely B-PROP) based on BERT for ad-hoc retrieval. The key idea is to use the powerful contextual language model BERT to replace the classical unigram language model for the ROP task construction, and re-train BERT itself towards the tailored objective for IR. Specifically, we introduce a novel contrastive method, inspired by the divergence-from-randomness idea, to leverage BERT's self-attention mechanism to sample representative words from the document. By further fine-tuning on downstream ad-hoc retrieval tasks, our method achieves significant improvements over baselines without pre-training or with other pre-training methods, and further pushes forward the SOTA on a variety of ad-hoc retrieval tasks.
Accepted by SIGIR 2021
References in corpus (6)
- Language Models are Few-Shot Learners
- A Deep Relevance Matching Model for Ad-hoc Retrieval
- Deeper Text Understanding for IR with Contextual Neural Language Modeling
- Simple Applications of BERT for Ad Hoc Document Retrieval
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Pre-training Tasks for Embedding-based Large-scale Retrieval
Cited by in corpus (7)
- Semantic Models for the First-stage Retrieval: A Comprehensive Review
- CorpusBrain: Pre-train a Generative Retrieval Model for Knowledge-Intensive Language Tasks
- Continual Learning for Generative Retrieval over Dynamic Corpora
- Pre-train a Discriminative Text Encoder for Dense Retrieval via Contrastive Span Prediction
- Certified Robustness to Word Substitution Ranking Attack for Neural Ranking Models
- Pre-training for Ad-hoc Retrieval: Hyperlink is Also You Need
- YES SIR!Optimizing Semantic Space of Negatives with Self-Involvement Ranker