2 citations · 2 across the 2 of their papers we have counts for
11 papers
MSA-UNet3+: Multi-Scale Attention UNet3+ with New Supervised Prototypical Contrastive Loss for Coronary DSA Image Segmentation
Rayan Merghani Ahmed, Adnan Iltaf, Mohamed Elmanna +5
Accurate segmentation of coronary Digital Subtraction Angiography (DSA) images is essential for diagnosing and treating coronary artery disease (CAD). Despite advances in deep lear…
SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards
Sheng Xia, Zhengqin Lai, Tianxiang Jiang +4
Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-tempo…
DV-VLN: Dual Verification for Reliable LLM-Based Vision-and-Language Navigation
Zijun Li, Shijie Li, Zhenxi Zhang +2
Vision-and-Language Navigation (VLN) requires an embodied agent to navigate in a complex 3D environment according to natural language instructions. Recent progress in large languag…
ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images
Hongyu Ge, Longkun Hao, Zihui Xu +5
Medical Visual Question Answering (Med-VQA) represents a critical and challenging subtask within the general VQA domain. Despite significant advancements in general VQA, multimodal…
M-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
Shenxi Liu, Kan Li, Mingyang Zhao +5
With the rapid progress of artificial intelligence (AI) in multi-modal understanding, there is increasing potential for video comprehension technologies to support professional dom…
VesselSAM: Leveraging SAM for Aortic Vessel Segmentation with AtrousLoRA
Adnan Iltaf, Rayan Merghani Ahmed, Zhenxi Zhang +2
Medical image segmentation is crucial for clinical diagnosis and treatment planning, especially when dealing with complex anatomical structures such as vessels. However, accurately…