1 paper
Pengyu Wang, Baochen Xiong, Xiaoshan Yang +4
Vision-Language Models (VLMs) have demonstrated strong performance in multimodal understanding and generation. However, fine-tuning of VLMs typically relies on centralized data, wh…