24 papers
MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models
Yuan Wang, Hualiang Wang, Yixin Chen +6
Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches ei…
FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening
Bin Pu, Jiewen Yang, Liwen Wang +9
A large number of infants with congenital anomalies are born each year globally, especially in areas with underdeveloped medical resources. Currently, fetal ultrasound screening is…
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Hangjie Yuan, Yichen Qian, Zhiwei Tang +21
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric chall…
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
Zihan Xu, Yanzhen Chen, Xiaocheng Zhang +4
The paper presents LongMedBench, a benchmark built from MIMIC-IV electronic health records that evaluates medical agents on long-horizon clinical decision-making across multiple vi…
MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding
Yuan Wang, Shujian Gao, Songtao Jiang +2
Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right time. In real clinical settings,…
AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation
Yuan Wang, Wanxing Chang, Songtao Jiang +8
Traditional metrics for Medical Report Generation (MRG) predominantly rely on surface-level n-gram overlap, which fails to capture clinical factual accuracy and often overlooks cat…