99 citations · 208 across the 17 of their papers we have counts for
Showing 2023Show all
3 papers · 1 filter
cs.MM2023★ 1 cited
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
Chaoya Jiang, Haiyang Xu, Wei Ye +7
Vision-Language Pre-training (VLP) methods based on object detection enjoy the rich knowledge of fine-grained object-text alignment but at the cost of computationally expensive inf…
cs.CV2023
BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization
Chaoya Jiang, Haiyang Xu, Wei Ye +7
Vision Transformer (ViT) based Vision-Language Pre-training (VLP) models have demonstrated impressive performance in various tasks. However, the lengthy visual token sequences fed…
cs.CV2023★ 50 cited
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video
Haiyang Xu, Qinghao Ye, Ming Yan +12
Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for…