1 paper
Chengxin Liu, Wonseok Choi, Chenshuang Zhang +1
Vision-Language Models (VLMs) have demonstrated strong capability in a wide range of tasks such as visual recognition, document parsing, and visual grounding. Nevertheless, recent…