20 citations · 43 across the 4 of their papers we have counts for
6 papers
Grid-VLP: Revisiting Grid Features for Vision-Language Pre-training
Ming Yan, Haiyang Xu, Chenliang Li +4
Existing approaches to vision-language pre-training (VLP) heavily rely on an object detector based on bounding boxes (regions), where salient objects are first detected from images…
E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning
Haiyang Xu, Ming Yan, Chenliang Li +4
Vision-language pre-training (VLP) on large-scale image-text pairs has achieved huge success for the cross-modal downstream tasks. The most existing pre-training methods mainly ado…
StructuralLM: Structural Pre-training for Form Understanding
Chenliang Li, Bin Bi, Ming Yan +4
Large pre-trained language models achieve state-of-the-art results when fine-tuned on downstream NLP tasks. However, they almost exclusively focus on text-only representation, whil…
SemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels
Chenliang Li, Ming Yan, Haiyang Xu +4
Vision-language pre-training (VLP) on large-scale image-text pairs has recently witnessed rapid progress for learning cross-modal representations. Existing pre-training methods eit…
PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned Generation
Bin Bi, Chenliang Li, Chen Wu +5
Self-supervised pre-training, such as BERT, MASS and BART, has emerged as a powerful technique for natural language understanding and generation. Existing pre-training techniques e…
Incorporating External Knowledge into Machine Reading for Generative Question Answering
Bin Bi, Chen Wu, Ming Yan +3
Commonsense and background knowledge is required for a QA model to answer many nontrivial questions. Different from existing work on knowledge-aware QA, we focus on a more challeng…