1 paper
Mengmeng Zhang, Xiaoping Wu, Hao Luo +2
Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We posit this limitation arises from the sca…