1 paper · 1 filter
Yifan Wang, Hongfeng Ai, Quangao Liu +7
Vision Language Models (VLMs) face challenges in effectively coordinating diverse attention mechanisms for cross-modal embedding learning, leading to mismatched attention and subop…