575 citations · 1k across the 20 of their papers we have counts for
12 papers · 1 filter
TIAGE: A Benchmark for Topic-Shift Aware Dialog Modeling
Huiyuan Xie, Zhenghao Liu, Chenyan Xiong +2
Human conversations naturally evolve around different topics and fluently move between them. In research on dialog systems, the ability to actively and smoothly transition to new t…
Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval
Chen Zhao, Chenyan Xiong, Jordan Boyd-Graber +1
Complex question answering often requires finding a reasoning chain that consists of multiple evidence pieces. Current approaches incorporate the strengths of structured knowledge…
Data Augmentation for Abstractive Query-Focused Multi-Document Summarization
Ramakanth Pasunuru, Asli Celikyilmaz, Michel Galley +4
The progress in Query-focused Multi-Document Summarization (QMDS) has been limited by the lack of sufficient largescale high-quality training datasets. We present two QMDS training…
COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining
Yu Meng, Chenyan Xiong, Payal Bajaj +4
We present a self-supervised learning framework, COCO-LM, that pretrains Language Models by COrrecting and COntrasting corrupted text sequences. Following ELECTRA-style pretraining…
Text Classification Using Label Names Only: A Language Model Self-Training Approach
Yu Meng, Yunyi Zhang, Jiaxin Huang +4
Current text classification methods typically require a good number of human-labeled documents as training data, which can be costly and difficult to obtain in real applications. H…
Knowledge-Aware Language Model Pretraining
Corby Rosset, Chenyan Xiong, Minh Phan +3
How much knowledge do pretrained language models hold? Recent research observed that pretrained transformers are adept at modeling semantics but it is unclear to what degree they g…