From the 1 of 20 linked papers with an AI index.
1 paper · 1 filter
Xunguang Wang, Zhenlan Ji, Pingchuan Ma +2
Large vision-language models (LVLMs) have demonstrated their incredible capability in image understanding and response generation. However, this rich visual interaction also makes…