activity
20192022
most citedSparTerm: Learning Term-based Sparse Representation for Fast Text Retrieval

59 citations · 79 across the 9 of their papers we have counts for

collaborators
Showing 2022Show all

6 papers · 1 filter

cs.CL2022

Retrieval-based Disentangled Representation Learning with Natural Language Supervision

Jiawei Zhou, Xiaoguang Li, Lifeng Shang +3

Disentangled representation learning remains challenging as the underlying factors of variation in the data do not naturally exist. The inherent complexity of real-world data makes…

cs.CL2022

Pre-training Language Models with Deterministic Factual Knowledge

Shaobo Li, Xiaoguang Li, Lifeng Shang +5

Previous works show that Pre-trained Language Models (PLMs) can capture factual knowledge. However, some analyses reveal that PLMs fail to perform it robustly, e.g., being sensitiv…

cs.IR2022★ 12 cited

Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages

Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo +6

MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval…

cs.CL2022

Hyperlink-induced Pre-training for Passage Retrieval in Open-domain Question Answering

Jiawei Zhou, Xiaoguang Li, Lifeng Shang +10

To alleviate the data scarcity problem in training question answering systems, recent works propose additional intermediate pre-training for dense passage retrieval (DPR). However,…

cs.CL2022

How Pre-trained Language Models Capture Factual Knowledge? A Causal-Inspired Analysis

Shaobo Li, Xiaoguang Li, Lifeng Shang +6

Recently, there has been a trend to investigate the factual knowledge captured by Pre-trained Language Models (PLMs). Many works show the PLMs' ability to fill in the missing factu…

cs.CL2022★ 3 cited

Read before Generate! Faithful Long Form Question Answering with Machine Reading

Dan Su, Xiaoguang Li, Jindi Zhang +4

Long-form question answering (LFQA) aims to generate a paragraph-length answer for a given question. While current work on LFQA using large pre-trained model for generation are eff…