1 paper · 1 filter
Hang Hua, Yunlong Tang, Ziyun Zeng +5
The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal understanding, enabling more sophisticated and accurate integration of visual and textual in…