2 citations · 2 across the 11 of their papers we have counts for
13 papers
A Vision-Language-Action Model for Adaptive Ultrasound-Guided Needle Insertion and Needle Tracking
Yuelin Zhang, Qingpeng Ding, Longxiang Tang +2
Ultrasound (US)-guided needle insertion is a critical yet challenging procedure due to dynamic imaging conditions and difficulties in needle visualization. Many methods have been p…
Beyond Ground-Truth: Leveraging Image Quality Priors for Real-World Image Restoration
Fengyang Xiao, Peng Hu, Lei Xu +7
Real-world image restoration aims to restore high-quality (HQ) images from degraded low-quality (LQ) inputs captured under uncontrolled conditions. Existing methods typically depen…
Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models
Chengyu Fang, Heng Guo, Zheng Jiang +3
Multimodal large language models are promising for clinical visual question answering tasks, but scaling to 3D imaging is hindered by high computational costs. Prior methods often…
Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning
Chubin Chen, Sujie Hu, Jiashu Zhu +8
Recent studies have demonstrated significant progress in aligning text-to-image diffusion models with human preference via Reinforcement Learning from Human Feedback. However, whil…
M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision
Che Liu, Zheng Jiang, Chengyu Fang +5
Medical image retrieval is essential for clinical decision-making and translational research, relying on discriminative visual representations. Yet, current methods remain fragment…
Reversible Unfolding Network for Concealed Visual Perception with Generative Refinement
Chunming He, Fengyang Xiao, Rihan Zhang +3
Existing methods for concealed visual perception (CVP) often leverage reversible strategies to decrease uncertainty, yet these are typically confined to the mask domain, leaving th…