activity
20222024
most citedmPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

50 citations · 120 across the 20 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV20242 cited

TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning

Liang Zhang, Anwen Hu, Haiyang Xu +5

Charts are important for presenting and explaining complex data relationships. Recently, multimodal large language models (MLLMs) have shown remarkable capabilities in various char…

cs.CV2024

Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval

Haowei Liu, Yaya Shi, Haiyang Xu +8

In video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representation…

cs.CV20235 cited

UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Jiabo Ye, Anwen Hu, Haiyang Xu +11

Text is ubiquitous in our visual world, conveying crucial information, such as in documents, websites, and everyday photographs. In this work, we propose UReader, a first explorati…

cs.CV20234 cited

Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks

Haiyang Xu, Qinghao Ye, Xuan Wu +13

To promote the development of Vision-Language Pre-training (VLP) and multimodal Large Language Model (LLM) in the Chinese community, we firstly release the largest public Chinese h…

cs.CV20232 cited

CIMI4D: A Large Multimodal Climbing Motion Dataset under Human-scene Interactions

Ming Yan, Xin Wang, Yudi Dai +5

Motion capture is a long-standing research problem. Although it has been studied for decades, the majority of research focus on ground-based movements such as walking, sitting, dan…

cs.CV202350 cited

mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Haiyang Xu, Qinghao Ye, Ming Yan +12

Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for…