collaborators

5 papers

cs.CV2026

Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training

Yingsheng Liu, Haiming Li, Jingmin Zhu +6

While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent in structured clinical tables.…

eess.IV2026

M2Diff: Multi-Modality Multi-Task Enhanced Diffusion Model for MRI-Guided Low-Dose PET Enhancement

Ghulam Nabi Ahmad Hassan Yar, Himashi Peiris, Victoria Mar +2

Positron emission tomography (PET) scans expose patients to radiation, which can be mitigated by reducing the dose, albeit at the cost of diminished quality. This makes low-dose (L…

cs.CV2026

A Vision-Language Foundation Model for Zero-shot Clinical Collaboration and Automated Concept Discovery in Dermatology

Siyuan Yan, Xieji Li, Dan Mo +28

Medical foundation models have shown promise in controlled benchmarks, yet widespread deployment remains hindered by reliance on task-specific fine-tuning. Here, we introduce DermF…

cs.CV2025

Multi-Aspect Knowledge-Enhanced Medical Vision-Language Pretraining with Multi-Agent Data Generation

Xieji Li, Siyuan Yan, Yingsheng Liu +4

Vision-language pretraining (VLP) has emerged as a powerful paradigm in medical image analysis, enabling representation learning from large-scale image-text pairs without relying o…

cs.CV2025

A Multimodal Vision Foundation Model for Clinical Dermatology

Siyuan Yan, Zhen Yu, Clare Primiero +22

Diagnosing and treating skin diseases require advanced visual skills across domains and the ability to synthesize information from multiple imaging modalities. While current deep l…