Neural Rankers for Effective Screening Prioritisation in Medical Systematic Review Literature Search
arXiv:2212.09017 · doi:10.1145/3572960.3572980
Abstract
Medical systematic reviews typically require assessing all the documents retrieved by a search. The reason is two-fold: the task aims for ``total recall''; and documents retrieved using Boolean search are an unordered set, and thus it is unclear how an assessor could examine only a subset. Screening prioritisation is the process of ranking the (unordered) set of retrieved documents, allowing assessors to begin the downstream processes of the systematic review creation earlier, leading to earlier completion of the review, or even avoiding screening documents ranked least relevant. Screening prioritisation requires highly effective ranking methods. Pre-trained language models are state-of-the-art on many IR tasks but have yet to be applied to systematic review screening prioritisation. In this paper, we apply several pre-trained language models to the systematic review document ranking task, both directly and fine-tuned. An empirical analysis compares how effective neural methods compare to traditional methods for this task. We also investigate different types of document representations for neural methods and their impact on ranking performance. Our results show that BERT-based rankers outperform the current state-of-the-art screening prioritisation methods. However, BERT rankers and existing methods can actually be complementary, and thus, further improvements may be achieved if used in conjunction.
References in corpus (5)
- Multi-Stage Document Ranking with BERT
- Pseudo-Relevance Feedback for Multiple Representation Dense Retrieval
- From Little Things Big Things Grow: A Collection with Seed Studies for Medical Systematic Review Literature Search
- To Interpolate or not to Interpolate: PRF, Dense and Sparse Retrievers
- Goldilocks: Just-Right Tuning of BERT for Technology-Assisted Review
Cited by in corpus (5)
- A Reproducibility and Generalizability Study of Large Language Models for Query Generation
- Stopping Methods for Technology Assisted Reviews based on Point Processes
- Outcome-based Evaluation of Systematic Review Automation
- Dense Retrieval with Continuous Explicit Feedback for Systematic Review Screening Prioritisation
- Reproducible Hybrid Time-Travel Retrieval in Evolving Corpora