activity
20192022
most citedERNIE: Enhanced Representation through Knowledge Integration

773 citations · 1k across the 10 of their papers we have counts for

collaborators

16 papers

cs.CV20223 cited

ERNIE-UniX2: A Unified Cross-lingual Cross-modal Framework for Understanding and Generation

Bin Shan, Yaqian Han, Weichong Yin +5

Recent cross-lingual cross-modal works attempt to extend Vision-Language Pre-training (VLP) models to non-English inputs and achieve impressive performance. However, these models f…

cs.CL20221 cited

Clip-Tuning: Towards Derivative-free Prompt Learning with a Mixture of Rewards

Yekun Chai, Shuohuan Wang, Yu Sun +3

Derivative-free prompt learning has emerged as a lightweight alternative to prompt tuning, which only requires model inference to optimize the prompts. However, existing work did n…

cs.CL20226 cited

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

Qiming Peng, Yinxu Pan, Wenjin Wang +12

Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and u…

cs.CV202210 cited

ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training

Bin Shan, Weichong Yin, Yu Sun +3

Recent Vision-Language Pre-trained (VLP) models based on dual encoder have attracted extensive attention from academia and industry due to their superior performance on various cro…

cs.CL202219 cited

ERNIE-Search: Bridging Cross-Encoder with Dual-Encoder via Self On-the-fly Distillation for Dense Passage Retrieval

Yuxiang Lu, Yiding Liu, Jiaxiang Liu +8

Neural retrievers based on pre-trained language models (PLMs), such as dual-encoders, have achieved promising performance on the task of open-domain question answering (QA). Their…

cs.CL20226 cited

ERNIE-SPARSE: Learning Hierarchical Efficient Transformer Through Regularized Self-Attention

Yang Liu, Jiaxiang Liu, Li Chen +7

Sparse Transformer has recently attracted a lot of attention since the ability for reducing the quadratic dependency on the sequence length. We argue that two factors, information…