6 citations · 6 across the 2 of their papers we have counts for
1 paper · 1 filter
Ziqi Wang, Chang Che, Qi Wang +3
Visual instruction tuning (VIT) enables multimodal large language models (MLLMs) to effectively handle a wide range of vision tasks by framing them as language-based instructions.…