1 paper
Yubo Zhang, Xueqing Wang, Manhui Lin +13
Vision-Language Models (VLMs) have achieved impressive results on general vision-language tasks, yet they suffer from hallucination, imprecise localization, and prohibitive computa…