activity
20162024
most citedERNIE: Enhanced Representation through Knowledge Integration

773 citations · 1.2k across the 35 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20223 cited

ERNIE-UniX2: A Unified Cross-lingual Cross-modal Framework for Understanding and Generation

Bin Shan, Yaqian Han, Weichong Yin +5

Recent cross-lingual cross-modal works attempt to extend Vision-Language Pre-training (VLP) models to non-English inputs and achieve impressive performance. However, these models f…

cs.CV20221 cited

CLOP: Video-and-Language Pre-Training with Knowledge Regularizations

Guohao Li, Hu Yang, Feng He +4

Video-and-language pre-training has shown promising results for learning generalizable representations. Most existing approaches usually model video and text in an implicit manner,…

cs.CV202210 cited

ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training

Bin Shan, Weichong Yin, Yu Sun +3

Recent Vision-Language Pre-trained (VLP) models based on dual encoder have attracted extensive attention from academia and industry due to their superior performance on various cro…

cs.CV20211 cited

Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal Matching

Bofeng Wu, Guocheng Niu, Jun Yu +3

This paper proposes an approach to Dense Video Captioning (DVC) without pairwise event-sentence annotation. First, we adopt the knowledge distilled from relevant and well solved ta…

cs.CV2020

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

Fei Yu, Jiji Tang, Weichong Yin +4

We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL…