1 paper
Ningxin Pan, Hanyu Li, Yehui Tang
Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such tasks, however, require more…