Passage Re-ranking with BERT
arXiv:1901.04085
Abstract
Recently, neural models pretrained on a language modeling task, such as ELMo (Peters et al., 2017), OpenAI GPT (Radford et al., 2018), and BERT (Devlin et al., 2018), have achieved impressive results on various natural language processing tasks such as question-answering and natural language inference. In this paper, we describe a simple re-implementation of BERT for query-based passage re-ranking. Our system is the state of the art on the TREC-CAR dataset and the top entry in the leaderboard of the MS MARCO passage retrieval task, outperforming the previous state of the art by 27% (relative) in MRR@10. The code to reproduce our results is available at https://github.com/nyu-dl/dl4marco-bert
References in corpus (2)
Cited by in corpus (154)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- End-to-End Open-Domain Question Answering with BERTserini
- Deeper Text Understanding for IR with Contextual Neural Language Modeling
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
- Document Expansion by Query Prediction
- Multi-Stage Document Ranking with BERT
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering
- Understanding the Behaviors of BERT in Ranking
- BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
- Simple Applications of BERT for Ad Hoc Document Retrieval
- Context-Aware Sentence/Passage Term Importance Estimation For First Stage Retrieval
- Semantic Models for the First-stage Retrieval: A Comprehensive Review
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Sparse, Dense, and Attentional Representations for Text Retrieval
- RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering
- Learning-to-Rank with BERT in TF-Ranking
- Complementing Lexical Retrieval with Semantic Residual Embedding
- RepBERT: Contextualized Text Embeddings for First-Stage Retrieval
- Distilling Dense Representations for Ranking using Tightly-Coupled Teachers
- Training Neural Response Selection for Task-Oriented Dialogue Systems
- A Transformer-based Embedding Model for Personalized Product Search
- Fast Passage Re-ranking with Contextualized Exact Term Matching and Efficient Passage Expansion
- An Objective Metric for Explainable AI: How and Why to Estimate the Degree of Explainability
- The Expando-Mono-Duo Design Pattern for Text Ranking with Pretrained Sequence-to-Sequence Models
- Distilling Knowledge from Reader to Retriever for Question Answering
- Adversarial Retriever-Ranker for dense text retrieval
- An Updated Duet Model for Passage Re-ranking
- ABNIRML: Analyzing the Behavior of Neural IR Models
- Learning a Product Relevance Model from Click-Through Data in E-Commerce
- BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
- Investigating the Successes and Failures of BERT for Passage Re-Ranking
- Mixed Attention Transformer for Leveraging Word-Level Knowledge to Neural Cross-Lingual Information Retrieval
- Is Retriever Merely an Approximator of Reader?
- Article Reranking by Memory-Enhanced Key Sentence Matching for Detecting Previously Fact-Checked Claims
- Embedding-based Zero-shot Retrieval through Query Generation
- Sensitive Data Detection and Classification in Spanish Clinical Text: Experiments with BERT
- Multi-step Entity-centric Information Retrieval for Multi-Hop Question Answering
- Neural Passage Retrieval with Improved Negative Contrast
- Incorporating Query Term Independence Assumption for Efficient Retrieval and Ranking using Deep Neural Networks
- Document Ranking with a Pretrained Sequence-to-Sequence Model
- Conformer-Kernel with Query Term Independence for Document Retrieval
- Blockwise Self-Attention for Long Document Understanding
- Passage Ranking with Weak Supervision
- B-PROP: Bootstrapped Pre-training with Representative Words Prediction for Ad-hoc Retrieval
- Leveraging Semantic and Lexical Matching to Improve the Recall of Document Retrieval Systems: A Hybrid Approach
- The Power of Selecting Key Blocks with Local Pre-ranking for Long Document Information Retrieval
- Learning To Retrieve: How to Train a Dense Retrieval Model Effectively and Efficiently
- Taming Pretrained Transformers for Extreme Multi-label Text Classification
- BoostingBERT:Integrating Multi-Class Boosting into BERT for NLP Tasks
- NaturalProofs: Mathematical Theorem Proving in Natural Language
- TED: A Pretrained Unsupervised Summarization Model with Theme Modeling and Denoising
- Domain-Specific Pretraining for Vertical Search: Case Study on Biomedical Literature
- Let's measure run time! Extending the IR replicability infrastructure to include performance aspects
- TU Wien @ TREC Deep Learning '19 -- Simple Contextualization for Re-ranking
- SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval
- Optimizing Dense Retrieval Model Training with Hard Negatives
- You Only Compress Once: Towards Effective and Elastic BERT Compression via Exploit-Explore Stochastic Nature Gradient
- COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List
- Intra-Document Cascading: Learning to Select Passages for Neural Document Ranking
- Beyond Domain APIs: Task-oriented Conversational Modeling with Unstructured Knowledge Access Track in DSTC9
- TREC 2020 Podcasts Track Overview
- Co-BERT: A Context-Aware BERT Retrieval Model Incorporating Local and Query-specific Context
- Multi-Stage Conversational Passage Retrieval: An Approach to Fusing Term Importance Estimation and Neural Query Rewriting
- Leveraging Lead Bias for Zero-shot Abstractive News Summarization
- Rethink Training of BERT Rerankers in Multi-Stage Retrieval Pipeline
- CMT in TREC-COVID Round 2: Mitigating the Generalization Gaps from Web to Special Domain Search
- Semantic Labeling Using a Deep Contextualized Language Model
- Few-Shot Generative Conversational Query Rewriting
- Guided Transformer: Leveraging Multiple External Sources for Representation Learning in Conversational Search
- Learning Passage Impacts for Inverted Indexes
- Beyond English-Only Reading Comprehension: Experiments in Zero-Shot Multilingual Transfer for Bulgarian
- Exploring Research Interest in Stack Overflow -- A Systematic Mapping Study and Quality Evaluation
- Simplified TinyBERT: Knowledge Distillation for Document Retrieval
- Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval
- Zero-shot Neural Passage Retrieval via Domain-targeted Synthetic Question Generation
- Understanding the Effectiveness of Reviews in E-commerce Top-N Recommendation
- Transformer Based Language Models for Similar Text Retrieval and Ranking
- Cross-Lingual Relevance Transfer for Document Retrieval
- Longformer for MS MARCO Document Re-ranking Task
- The False COVID-19 Narratives That Keep Being Debunked: A Spatiotemporal Analysis
- Robust Layout-aware IE for Visually Rich Documents with Pre-trained Language Models
- Searching Scientific Literature for Answers on COVID-19 Questions
- Model Extraction and Defenses on Generative Adversarial Networks
- Multistage BiCross encoder for multilingual access to COVID-19 health information
- Delaying Interaction Layers in Transformer-based Encoders for Efficient Open Domain Question Answering
- BERT-QE: Contextualized Query Expansion for Document Re-ranking
- Pseudo Relevance Feedback with Deep Language Models and Dense Retrievers: Successes and Pitfalls
- Revisiting the Open-Domain Question Answering Pipeline
- PGT: Pseudo Relevance Feedback Using a Graph-Based Transformer
- Neural document expansion for ad-hoc information retrieval
- Discovering Useful Sentence Representations from Large Pretrained Language Models
- An Optimal Algorithm for Finding Champions in Tournament Graphs
- Brown University at TREC Deep Learning 2019
- Fairness Through Regularization for Learning to Rank
- IntenT5: Search Result Diversification using Causal Language Models
- Speech Recognition by Simply Fine-tuning BERT
- Curriculum Learning Strategies for IR: An Empirical Study on Conversation Response Ranking
- Contextualized Query Embeddings for Conversational Search
- DoSSIER@COLIEE 2021: Leveraging dense retrieval and summarization-based re-ranking for case law retrieval
- Patient Cohort Retrieval using Transformer Language Models
- Modeling Relevance Ranking under the Pre-training and Fine-tuning Paradigm
- Linguistically Informed Masking for Representation Learning in the Patent Domain
- Conversational Answer Generation and Factuality for Reading Comprehension Question-Answering
- SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking
- Long Document Ranking with Query-Directed Sparse Transformer
- Context-based Transformer Models for Answer Sentence Selection
- On the Calibration and Uncertainty of Neural Learning to Rank Models
- Cross-Lingual Training with Dense Retrieval for Document Retrieval
- On the Importance of Adaptive Data Collection for Extremely Imbalanced Pairwise Tasks
- BERTnesia: Investigating the capture and forgetting of knowledge in BERT
- Towards Confident Machine Reading Comprehension
- Dealing with Typos for BERT-based Passage Retrieval and Ranking
- GLOW : Global Weighted Self-Attention Network for Web Search
- Knowledge-driven Answer Generation for Conversational Search
- Weakly-Supervised Neural Response Selection from an Ensemble of Task-Specialised Dialogue Agents
- Multi-Perspective Semantic Information Retrieval in the Biomedical Domain
- RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking
- DeText: A Deep Text Ranking Framework with BERT
- Optimize What You Evaluate With: A Simple Yet Effective Framework For Direct Optimization Of IR Metrics
- Building an Efficient and Effective Retrieval-based Dialogue System via Mutual Learning
- An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking
- Pre-trained Language Model based Ranking in Baidu Search
- Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins
- Pre-training for Ad-hoc Retrieval: Hyperlink is Also You Need
- Duet at TREC 2019 Deep Learning Track
- Few-Shot Text Ranking with Meta Adapted Synthetic Weak Supervision
- Learning Better Sentence Representation with Syntax Information
- Training Adaptive Computation for Open-Domain Question Answering with Computational Constraints
- More Robust Dense Retrieval with Contrastive Dual Learning
- Combining Lexical and Dense Retrieval for Computationally Efficient Multi-hop Question Answering
- Exploiting Sentence-Level Representations for Passage Ranking
- Quality and Cost Trade-offs in Passage Re-ranking Task
- Cross-language Information Retrieval
- Deep Natural Language Processing for LinkedIn Search
- Dual Reader-Parser on Hybrid Textual and Tabular Evidence for Open Domain Question Answering
- MS MARCO: Benchmarking Ranking Models in the Large-Data Regime
- Text-to-Text Multi-view Learning for Passage Re-ranking
- BERT for Target Apps Selection: Analyzing the Diversity and Performance of BERT in Unified Mobile Search
- Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval
- Explaining Documents' Relevance to Search Queries
- A Data-Centric Framework for Composable NLP Workflows
- Autoregressive Reasoning over Chains of Facts with Transformers
- Learning to Expand: Reinforced Pseudo-relevance Feedback Selection for Information-seeking Conversations
- Mitigating the Position Bias of Transformer Models in Passage Re-Ranking
- Beyond [CLS] through Ranking by Generation
- Leveraging Query Resolution and Reading Comprehension for Conversational Passage Retrieval
- Navigation-Based Candidate Expansion and Pretrained Language Models for Citation Recommendation
- BERT Embeddings Can Track Context in Conversational Search
- Topic Propagation in Conversational Search
- Modularized Transfomer-based Ranking Framework
- Multi-Perspective Semantic Information Retrieval
- RadLex Normalization in Radiology Reports
- Using the Hammer Only on Nails: A Hybrid Method for Evidence Retrieval for Question Answering