activity
20242026
most citedHulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

1 citations · 1 across the 17 of their papers we have counts for

collaborators

19 papers

cs.CL2026

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

Zijie Meng, Yuncheng Zhang, Hualiang Wang +8

Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the m…

cs.CV2026

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

Yuan Wang, Hualiang Wang, Yixin Chen +6

Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches ei…

cs.CV2026

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

Yuan Wang, Shujian Gao, Songtao Jiang +2

Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right time. In real clinical settings,…

cs.CE2026

AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation

Yuan Wang, Wanxing Chang, Songtao Jiang +8

Traditional metrics for Medical Report Generation (MRG) predominantly rely on surface-level n-gram overlap, which fails to capture clinical factual accuracy and often overlooks cat…

cs.CV2026

How Far Are Video Models from True Multimodal Reasoning?

Xiaotian Zhang, Jianhui Wei, Yuan Wang +9

Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true multimodal reasoning? Existing…

cs.CV2026

Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos

Songtao Jiang, Sibo Song, Chenyi Zhou +12

The transition from image to video understanding requires vision-language models (VLMs) to shift from recognizing static patterns to reasoning over temporal dynamics such as motion…