1 citations · 1 across the 17 of their papers we have counts for
19 papers
MedUAG: Unified Understanding and Generation for Medical Multimodal Models
Zijie Meng, Yuncheng Zhang, Hualiang Wang +8
Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the m…
MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models
Yuan Wang, Hualiang Wang, Yixin Chen +6
Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches ei…
MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding
Yuan Wang, Shujian Gao, Songtao Jiang +2
Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right time. In real clinical settings,…
AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation
Yuan Wang, Wanxing Chang, Songtao Jiang +8
Traditional metrics for Medical Report Generation (MRG) predominantly rely on surface-level n-gram overlap, which fails to capture clinical factual accuracy and often overlooks cat…
How Far Are Video Models from True Multimodal Reasoning?
Xiaotian Zhang, Jianhui Wei, Yuan Wang +9
Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true multimodal reasoning? Existing…
Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos
Songtao Jiang, Sibo Song, Chenyi Zhou +12
The transition from image to video understanding requires vision-language models (VLMs) to shift from recognizing static patterns to reasoning over temporal dynamics such as motion…