1 paper · 1 filter
Yichen Yan, Ming Zhong, Qi Zhu +3
Multimodal large language models (MLLMs) rely heavily on instruction tuning to align vision and language capabilities, yet the computational cost of training on large-scale dataset…