activity
20232025
most citedmPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

29 citations · 35 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC

Haowei Liu, Xi Zhang, Haiyang Xu +8

In the field of MLLM-based GUI agents, compared to smartphones, the PC scenario not only features a more complex interactive environment, but also involves more intricate intra- an…

cs.CV20246 cited

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Jiabo Ye, Haiyang Xu, Haowei Liu +6

Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in executing instructions for a variety of single-image tasks. Despite this progress, significan…

cs.CV20241 cited

MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Haowei Liu, Xi Zhang, Haiyang Xu +8

Built on the power of LLMs, numerous multimodal large language models (MLLMs) have recently achieved remarkable performance on various vision-language tasks. However, most existing…

cs.CV2024

Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training

Haowei Liu, Yaya Shi, Haiyang Xu +8

In vision-language pre-training (VLP), masked image modeling (MIM) has recently been introduced for fine-grained cross-modal alignment. However, in most existing methods, the recon…

cs.CV2024

Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval

Haowei Liu, Yaya Shi, Haiyang Xu +8

In video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representation…