7 citations · 13 across the 4 of their papers we have counts for
4 papers
Global and Local Semantic Completion Learning for Vision-Language Pre-training
Rong-Cheng Tu, Yatai Ji, Jie Jiang +6
Cross-modal alignment plays a crucial role in vision-language pre-training (VLP) models, enabling them to capture meaningful associations across different modalities. For this purp…
Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition
Xinzhe Ni, Yong Liu, Hao Wen +3
Current methods for few-shot action recognition mainly fall into the metric learning framework following ProtoNet, which demonstrates the importance of prototypes. Although they ac…
Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning
Yatai Ji, Rongcheng Tu, Jie Jiang +6
Cross-modal alignment is essential for vision-language pre-training (VLP) models to learn the correct corresponding information across different modalities. For this purpose, inspi…
MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model
Yatai Ji, Junjie Wang, Yuan Gong +6
Multimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets. Such uncertainty is problematic for our i…