1 paper
Yuhao Su, Anwesa Choudhuri, Zhongpai Gao +8
Large vision-language models struggle with medical video understanding, where spatial precision, temporal reasoning, and clinical semantics are critical. To address this, we first…