9 papers
MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models
Yuan Wang, Hualiang Wang, Yixin Chen +6
Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches ei…
Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
Xiaotian Zhang, Yuan Wang, Ruizhe Chen +3
The deployment of Large Language Models (LLMs) in interactive systems necessitates a deep alignment with the nuanced and dynamic preferences of individual users. Current alignment…
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Songtao Jiang, Yuan Wang, Sibo Song +22
Real-world clinical decision-making requires integrating heterogeneous data, including medical text, 2D images, 3D volumes, and videos, while existing AI systems fail to unify all…
Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset
Ziye Deng, Ruihan He, Jiaxiang Liu +5
Medical image grounding aims to align natural language phrases with specific regions in medical images, serving as a foundational task for intelligent diagnosis, visual question an…
Modest-Align: Data-Efficient Alignment for Vision-Language Models
Jiaxiang Liu, Yuan Wang, Jiawei Du +3
Cross-modal alignment aims to map heterogeneous modalities into a shared latent space, as exemplified by models like CLIP, which benefit from large-scale image-text pretraining for…
V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
Yuan Wang, Jiaxiang Liu, Shujian Gao +5
Recent advances in multimodal techniques have led to significant progress in Medical Visual Question Answering (Med-VQA). However, most existing models focus on global image featur…