59 citations · 79 across the 8 of their papers we have counts for
8 papers
Pre-training Language Models with Deterministic Factual Knowledge
Shaobo Li, Xiaoguang Li, Lifeng Shang +5
Previous works show that Pre-trained Language Models (PLMs) can capture factual knowledge. However, some analyses reveal that PLMs fail to perform it robustly, e.g., being sensitiv…
Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages
Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo +6
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval…
Hyperlink-induced Pre-training for Passage Retrieval in Open-domain Question Answering
Jiawei Zhou, Xiaoguang Li, Lifeng Shang +10
To alleviate the data scarcity problem in training question answering systems, recent works propose additional intermediate pre-training for dense passage retrieval (DPR). However,…
How Pre-trained Language Models Capture Factual Knowledge? A Causal-Inspired Analysis
Shaobo Li, Xiaoguang Li, Lifeng Shang +6
Recently, there has been a trend to investigate the factual knowledge captured by Pre-trained Language Models (PLMs). Many works show the PLMs' ability to fill in the missing factu…
Read before Generate! Faithful Long Form Question Answering with Machine Reading
Dan Su, Xiaoguang Li, Jindi Zhang +4
Long-form question answering (LFQA) aims to generate a paragraph-length answer for a given question. While current work on LFQA using large pre-trained model for generation are eff…
Unsupervised Open-Domain Question Answering
Pengfei Zhu, Xiaoguang Li, Jian Li +1
Open-domain Question Answering (ODQA) has achieved significant results in terms of supervised learning manner. However, data annotation cannot also be irresistible for its huge dem…