10 papers
LocAnyMed: Vision-Language Grounding for Multimodal Medical Images
Zihan Wang, Tong Liu, Zhiwei Wang +6
Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. H…
Echo-DM: Ultrasound Marker Removal via Conditional Latent Diffusion and Region-Aware Fusion
Zhiwei Wang, Tao Huang, Wentao Jiang +7
Clinical ultrasound images often contain artificial markers, such as measurement calipers and text, to assist diagnostic interpretation and comparison. However, these markers can i…
SAMe: A Semantic Anatomy Mapping Engine for Robotic Ultrasound
Jing Zhang, Duojie Chen, Wentao Jiang +7
Robotic ultrasound has advanced local image-driven control, contact regulation, and view optimization, yet current systems lack the anatomical understanding needed to determine wha…
Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation
Jing Zhang, Wentao Jiang, Tao Huang +8
Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of these capabilities: special…
SARMAE: Masked Autoencoder for SAR Representation Learning
Danxu Liu, Di Wang, Hebaixu Wang +6
Synthetic Aperture Radar (SAR) imagery plays a critical role in all-weather, day-and-night remote sensing applications. However, existing SAR-oriented deep learning is constrained…
GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
Di Wang, Shunyu Liu, Wentao Jiang +10
Multimodal large language models (MLLMs) have undergone rapid development in advancing geospatial scene understanding. Recent studies have sought to enhance the reasoning capabilit…