collaborators

6 papers

cs.CV2026

LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

Zihan Wang, Tong Liu, Zhiwei Wang +6

Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. H…

cs.CV2026

Echo-DM: Ultrasound Marker Removal via Conditional Latent Diffusion and Region-Aware Fusion

Zhiwei Wang, Tao Huang, Wentao Jiang +7

Clinical ultrasound images often contain artificial markers, such as measurement calipers and text, to assist diagnostic interpretation and comparison. However, these markers can i…

cs.CV2026

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

Yuancheng Wei, Haojie Zhang, Linli Yao +7

Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key benchmark for fine-grained change…

cs.CV2026

Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation

Jing Zhang, Wentao Jiang, Tao Huang +8

Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of these capabilities: special…

cs.CV2025

Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward

Muhammad Islam, Tao Huang, Euijoon Ahn +1

This paper presents an in-depth survey on the use of multimodal Generative Artificial Intelligence (GenAI) and autoregressive Large Language Models (LLMs) for human motion understa…

eess.IV2025

Computationally Efficient Diffusion Models in Medical Imaging: A Comprehensive Review

Abdullah, Tao Huang, Ickjai Lee +1

The diffusion model has recently emerged as a potent approach in computer vision, demonstrating remarkable performances in the field of generative artificial intelligence. Capable…