172 citations · 226 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 4 cited
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks
Haiyang Xu, Qinghao Ye, Xuan Wu +13
To promote the development of Vision-Language Pre-training (VLP) and multimodal Large Language Model (LLM) in the Chinese community, we firstly release the largest public Chinese h…
cs.CV2023★ 50 cited
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video
Haiyang Xu, Qinghao Ye, Ming Yan +12
Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for…
cs.IR2022★ 172 cited
Multi-Behavior Hypergraph-Enhanced Transformer for Sequential Recommendation
Yuhao Yang, Chao Huang, Lianghao Xia +3
Learning dynamic user preference has become an increasingly important component for many online platforms (e.g., video-sharing sites, e-commerce systems) to make sequential recomme…