collaborators

5 papers

cs.CV2026

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

Kaichen Li, Zhilin Zhu, Jianhao Huang +7

In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Exi…

cs.LG2026

UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity

Pengyu Wang, Baochen Xiong, Xiaoshan Yang +4

Vision-Language Models (VLMs) have demonstrated strong performance in multimodal understanding and generation. However, fine-tuning of VLMs typically relies on centralized data, wh…

cs.AI2026

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

Zibo Shao, Baochen Xiong, Chengdong Xu +6

Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are…

cs.CL2026

Decoupled Vision-Language System for Multimodal Understanding and Generation

Yifan Xu, Baochen Xiong, Xiaoshan Yang +3

We introduce a new architecture design for multimodal large language models (MLLMs), Libra, capable of both multimodal understanding and generation. Libra architecture contains one…

cs.AI2025

Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation

Yuyang Wanyan, Xi Zhang, Haiyang Xu +9

In recent years, Multimodal Large Language Models (MLLMs) have been extensively utilized for multimodal reasoning tasks, including Graphical User Interface (GUI) automation. Unlike…