works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.AI2026

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

Zibo Shao, Baochen Xiong, Chengdong Xu +6

Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are…

cs.CL2026

Decoupled Vision-Language System for Multimodal Understanding and Generation

Yifan Xu, Baochen Xiong, Xiaoshan Yang +3

We introduce a new architecture design for multimodal large language models (MLLMs), Libra, capable of both multimodal understanding and generation. Libra architecture contains one…

cs.CV2026

PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging

Zibo Shao, Baochen Xiong, Xiaoshan Yang +4

The paper introduces PivotMerge, a framework for merging multimodal language models trained on heterogeneous data by separating shared cross‑modal alignment patterns from domain‑sp…

cs.LG2026

A Step Toward Federated Pretraining of Multimodal Large Language Models

Baochen Xiong, Yifan Xu, Xiaoshan Yang +3

The rapid evolution of Multimodal Large Language Models (MLLMs) is bottlenecked by the saturation of high-quality public data, while vast amounts of diverse multimodal data remain…

cs.LG2025

Pilot: Building the Federated Multimodal Instruction Tuning Framework

Baochen Xiong, Xiaoshan Yang, Yaguang Song +2

In this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal…