From the 1 of 5 linked papers with an AI index.
5 papers
AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models
Zibo Shao, Baochen Xiong, Chengdong Xu +6
Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are…
Decoupled Vision-Language System for Multimodal Understanding and Generation
Yifan Xu, Baochen Xiong, Xiaoshan Yang +3
We introduce a new architecture design for multimodal large language models (MLLMs), Libra, capable of both multimodal understanding and generation. Libra architecture contains one…
PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging
Zibo Shao, Baochen Xiong, Xiaoshan Yang +4
The paper introduces PivotMerge, a framework for merging multimodal language models trained on heterogeneous data by separating shared cross‑modal alignment patterns from domain‑sp…
A Step Toward Federated Pretraining of Multimodal Large Language Models
Baochen Xiong, Yifan Xu, Xiaoshan Yang +3
The rapid evolution of Multimodal Large Language Models (MLLMs) is bottlenecked by the saturation of high-quality public data, while vast amounts of diverse multimodal data remain…
Pilot: Building the Federated Multimodal Instruction Tuning Framework
Baochen Xiong, Xiaoshan Yang, Yaguang Song +2
In this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal…