RepBERT: Contextualized Text Embeddings for First-Stage Retrieval
arXiv:2006.15498
Abstract
Although exact term match between queries and documents is the dominant method to perform first-stage retrieval, we propose a different approach, called RepBERT, to represent documents and queries with fixed-length contextualized embeddings. The inner products of query and document embeddings are regarded as relevance scores. On MS MARCO Passage Ranking task, RepBERT achieves state-of-the-art results among all initial retrieval techniques. And its efficiency is comparable to bag-of-words methods.
For corresponding code and data, see https://github.com/jingtaozhan/RepBERT-Index
References in corpus (4)
- Convolutional Neural Network Architectures for Matching Natural Language Sentences
- REALM: Retrieval-Augmented Language Model Pre-Training
- Context-Aware Sentence/Passage Term Importance Estimation For First Stage Retrieval
- Microsoft AI Challenge India 2018: Learning to Rank Passages for Web Question Answering with Deep Attention Networks
Cited by in corpus (9)
- Fast Passage Re-ranking with Contextualized Exact Term Matching and Efficient Passage Expansion
- Pre-train a Discriminative Text Encoder for Dense Retrieval via Contrastive Span Prediction
- A Replication Study of Dense Passage Retriever
- CharacterBERT and Self-Teaching for Improving the Robustness of Dense Retrievers on Queries with Typos
- To Interpolate or not to Interpolate: PRF, Dense and Sparse Retrievers
- Optimizing Dense Retrieval Model Training with Hard Negatives
- Asyncval: A Toolkit for Asynchronously Validating Dense Retriever Checkpoints during Training
- Jointly Optimizing Query Encoder and Product Quantization to Improve Retrieval Performance
- Dealing with Typos for BERT-based Passage Retrieval and Ranking