Multistage BiCross encoder for multilingual access to COVID-19 health information
arXiv:2101.03013 · doi:10.1371/journal.pone.0256874
Abstract
The Coronavirus (COVID-19) pandemic has led to a rapidly growing 'infodemic' of health information online. This has motivated the need for accurate semantic search and retrieval of reliable COVID-19 information across millions of documents, in multiple languages. To address this challenge, this paper proposes a novel high precision and high recall neural Multistage BiCross encoder approach. It is a sequential three-stage ranking pipeline which uses the Okapi BM25 retrieval algorithm and transformer-based bi-encoder and cross-encoder to effectively rank the documents with respect to the given query. We present experimental results from our participation in the Multilingual Information Access (MLIA) shared task on COVID-19 multilingual semantic search. The independently evaluated MLIA results validate our approach and demonstrate that it outperforms other state-of-the-art approaches according to nearly all evaluation metrics in cases of both monolingual and bilingual runs.
References in corpus (8)
- SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
- Simple Applications of BERT for Ad Hoc Document Retrieval
- Classification Aware Neural Topic Model and its Application on a New COVID-19 Disinformation Corpus
- PARADE: Passage Representation Aggregation for Document Reranking
- The Expando-Mono-Duo Design Pattern for Text Ranking with Pretrained Sequence-to-Sequence Models
- Let's measure run time! Extending the IR replicability infrastructure to include performance aspects
- Cross-Lingual Relevance Transfer for Document Retrieval