1 paper
Jinhe Bi, Aniri, Zengjie Jin +11
Visual instruction tuning adapts pre-trained Multimodal Large Language Models (MLLMs) to follow human instructions for real-world applications. However, the rapid growth of these d…