1 paper
Shilong Zhang, Peize Sun, Shoufa Chen +6
Visual instruction tuning large language model(LLM) on image-text pairs has achieved general-purpose vision-language abilities. However, the lack of region-text pairs limits their…