1 paper
Kangyu Zhu, Ziyuan Qin, Huahui Yi +4
While mainstream vision-language models (VLMs) have advanced rapidly in understanding image level information, they still lack the ability to focus on specific areas designated by…