206 citations · 488 across the 59 of their papers we have counts for
26 papers · 1 filter
Wukong-Reader: Multi-modal Pre-training for Fine-grained Visual Document Understanding
Haoli Bai, Zhiguang Liu, Xiaojun Meng +9
Unsupervised pre-training on millions of digital-born or scanned documents has shown promising advances in visual document understanding~(VDU). While various vision-language pre-tr…
KPT: Keyword-guided Pre-training for Grounded Dialog Generation
Qi Zhu, Fei Mi, Zheng Zhang +6
Incorporating external knowledge into the response generation process is essential to building more helpful and reliable dialog agents. However, collecting knowledge-grounded conve…
Retrieval-based Disentangled Representation Learning with Natural Language Supervision
Jiawei Zhou, Xiaoguang Li, Lifeng Shang +3
Disentangled representation learning remains challenging as the underlying factors of variation in the data do not naturally exist. The inherent complexity of real-world data makes…
G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
Zhongwei Wan, Yichun Yin, Wei Zhang +5
Recently, domain-specific PLMs have been proposed to boost the task performance of specific domains (e.g., biomedical and computer science) by continuing to pre-train general PLMs…
Lexicon-injected Semantic Parsing for Task-Oriented Dialog
Xiaojun Meng, Wenlin Dai, Yasheng Wang +4
Recently, semantic parsing using hierarchical representations for dialog systems has captured substantial attention. Task-Oriented Parse (TOP), a tree representation with intents a…
LiteVL: Efficient Video-Language Learning with Enhanced Spatial-Temporal Modeling
Dongsheng Chen, Chaofan Tao, Lu Hou +3
Recent large-scale video-language pre-trained models have shown appealing performance on various downstream tasks. However, the pre-training process is computationally expensive du…