16 papers
Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
Haoyang Chen, Jing Zhang, Hebaixu Wang +7
Multi-modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross-mod…
SAMe: A Semantic Anatomy Mapping Engine for Robotic Ultrasound
Jing Zhang, Duojie Chen, Wentao Jiang +7
Robotic ultrasound has advanced local image-driven control, contact regulation, and view optimization, yet current systems lack the anatomical understanding needed to determine wha…
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
Zhiming Luo, Di Wang, Haonan Guo +2
Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmarks remain heavily biased toward percep…
Residual Diffusion Bridge Model for Image Restoration
Hebaixu Wang, Jing Zhang, Haoyang Chen +4
Diffusion bridge models establish probabilistic paths between arbitrary paired distributions and exhibit great potential for universal image restoration. Most existing methods mere…
Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation
Jing Zhang, Wentao Jiang, Tao Huang +8
Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of these capabilities: special…
GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
Zixuan Song, Jing Zhang, Di Wang +5
Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradi…