1 paper · 1 filter
Massimo Bosetti, Shibingfeng Zhang, Benedetta Liberatori +3
Vision-language models (VLMs) have demonstrated remarkable performance across various visual tasks, leveraging joint learning of visual and textual representations. While these mod…