1 paper
Heng Zhang, Haichuan Hu, Yaomin Shen +9
Large Vision-Language Models (LVLMs) have demonstrated impressive performance on multimodal tasks through scaled architectures and extensive training. However, existing Mixture of…