most citedmPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

50 citations · 73 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20234 cited

Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks

Haiyang Xu, Qinghao Ye, Xuan Wu +13

To promote the development of Vision-Language Pre-training (VLP) and multimodal Large Language Model (LLM) in the Chinese community, we firstly release the largest public Chinese h…

cs.CL20232 cited

Distinguish Before Answer: Generating Contrastive Explanation as Knowledge for Commonsense Question Answering

Qianglong Chen, Guohai Xu, Ming Yan +4

Existing knowledge-enhanced methods have achieved remarkable results in certain QA tasks via obtaining diverse knowledge from different knowledge bases. However, limited by the pro…

cs.RO20231 cited

Active Velocity Estimation using Light Curtains via Self-Supervised Multi-Armed Bandits

Siddharth Ancha, Gaurav Pathak, Ji Zhang +2

To navigate in an environment safely and autonomously, robots must accurately estimate where obstacles are and how they move. Instead of using expensive traditional 3D sensors, we…

cs.CV202350 cited

mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Haiyang Xu, Qinghao Ye, Ming Yan +12

Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for…

cs.CV202216 cited

X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval

Yiwei Ma, Guohai Xu, Xiaoshuai Sun +3

Video-text retrieval has been a crucial and fundamental task in multi-modal research. The development of video-text retrieval has been considerably promoted by large-scale multi-mo…