4 papers
A Step Toward Federated Pretraining of Multimodal Large Language Models
Baochen Xiong, Yifan Xu, Xiaoshan Yang +3
The rapid evolution of Multimodal Large Language Models (MLLMs) is bottlenecked by the saturation of high-quality public data, while vast amounts of diverse multimodal data remain…
Harmony: A Unified Framework for Modality Incremental Learning
Yaguang Song, Xiaoshan Yang, Dongmei Jiang +2
Incremental learning aims to enable models to continuously acquire knowledge from evolving data streams while preserving previously learned capabilities. While current research pre…
Pilot: Building the Federated Multimodal Instruction Tuning Framework
Baochen Xiong, Xiaoshan Yang, Yaguang Song +2
In this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal…
Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC
Ming Tao, Bing-Kun Bao, Yaowei Wang +1
Large pretrained diffusion models have demonstrated impressive generation capabilities and have been adapted to various downstream tasks. However, unlike Large Language Models (LLM…