activity
20192025
most citedSemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels

20 citations · 61 across the 16 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2024

Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection

Wei Ye, Chaoya Jiang, Haiyang Xu +6

Vision Transformers (ViTs) have become increasingly popular in large-scale Vision and Language Pre-training (VLP) models. Although previous VLP research has demonstrated the effica…

cs.CV2024

Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training

Haowei Liu, Yaya Shi, Haiyang Xu +8

In vision-language pre-training (VLP), masked image modeling (MIM) has recently been introduced for fine-grained cross-modal alignment. However, in most existing methods, the recon…

cs.CV2024

Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval

Haowei Liu, Yaya Shi, Haiyang Xu +8

In video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representation…

cs.CV20235 cited

UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Jiabo Ye, Anwen Hu, Haiyang Xu +11

Text is ubiquitous in our visual world, conveying crucial information, such as in documents, websites, and everyday photographs. In this work, we propose UReader, a first explorati…

cs.CV2023

BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization

Chaoya Jiang, Haiyang Xu, Wei Ye +7

Vision Transformer (ViT) based Vision-Language Pre-training (VLP) models have demonstrated impressive performance in various tasks. However, the lengthy visual token sequences fed…

cs.CV20234 cited

Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks

Haiyang Xu, Qinghao Ye, Xuan Wu +13

To promote the development of Vision-Language Pre-training (VLP) and multimodal Large Language Model (LLM) in the Chinese community, we firstly release the largest public Chinese h…