most citedERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training

10 citations · 22 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV20223 cited

ERNIE-UniX2: A Unified Cross-lingual Cross-modal Framework for Understanding and Generation

Bin Shan, Yaqian Han, Weichong Yin +5

Recent cross-lingual cross-modal works attempt to extend Vision-Language Pre-training (VLP) models to non-English inputs and achieve impressive performance. However, these models f…

cs.CL20226 cited

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

Qiming Peng, Yinxu Pan, Wenjin Wang +12

Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and u…

cs.CV202210 cited

ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training

Bin Shan, Weichong Yin, Yu Sun +3

Recent Vision-Language Pre-trained (VLP) models based on dual encoder have attracted extensive attention from academia and industry due to their superior performance on various cro…

cs.CV20223 cited

ERNIE-mmLayout: Multi-grained MultiModal Transformer for Document Understanding

Wenjin Wang, Zhengjie Huang, Bin Luo +8

Recent efforts of multimodal Transformers have improved Visually Rich Document Understanding (VrDU) tasks via incorporating visual and textual information. However, existing approa…

cs.CV2020

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

Fei Yu, Jiji Tang, Weichong Yin +4

We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL…