16 citations · 27 across the 15 of their papers we have counts for
1 paper · 2 filters
Ruichuan An, Sihan Yang, Ming Lu +9
Current vision-language models (VLMs) show exceptional abilities across diverse tasks, such as visual question answering. To enhance user experience, recent studies investigate VLM…