4 papers
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
Yue Zhou, Erxuan Wu, Yikang Sun +5
Vision-Language Models (VLMs) have significantly advanced medical visual question answering, yet their performance in ultrasound remains suboptimal. In clinical practice, sonograph…
Motif-Video 2B: Technical Report
Junghwan Lim, Wai Ting Cheung, Minsu Ha +25
Training strong video generation models usually requires massive datasets, large parameter counts, and substantial compute. In this work, we ask whether strong text-to-video qualit…
From Adaptation to Generalization: Adaptive Visual Prompting for Medical Image Segmentation
Evren Ãetinkaya, Sangmin Lee, Jung Uk Kim +2
Visual prompting has emerged as a powerful method for adapting pre-trained models to new domains without updating model parameters. However, existing prompting methods typically op…
Leveraging Textual Compositional Reasoning for Robust Change Captioning
Kyu Ri Park, Jiyoung Park, Seong Tae Kim +2
Change captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful change…