1 paper · 1 filter
Shen Lin, Junhao Dong, Rongjie Chen +3
Vision-language models (VLMs) have shown remarkable ability in aligning visual and textual representations, enabling a wide range of multimodal applications. However, their large-s…