6 papers
HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts
Haozhe Luo, Ziyu Zhou, Shelley Zixin Shu +1
Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the dominant pretraining source.…
KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection
Haozhe Luo, Shelley Zixin Shu, Ziyu Zhou +2
Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and clinical text, thereby lever…
Lamps: Learning Anatomy from Multiple Perspectives via Self-supervision in Chest Radiographs
Ziyu Zhou, Haozhe Luo, Mohammad Reza Hosseinzadeh Taher +4
Foundation models have been successful in natural language processing and computer vision because they are capable of capturing the underlying structures (foundation) of natural la…
XBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography
Haozhe Luo, Shelley Zixin Shu, Ziyu Zhou +2
Vision-language models (VLMs) have recently shown remarkable zero-shot performance in medical image understanding, yet their grounding ability, the extent to which textual concepts…
RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
Chengrun Li, Corentin Royer, Haozhe Luo +6
Most current medical vision language models struggle to jointly generate diagnostic text and pixel-level segmentation masks in response to complex visual questions. This represents…
ACE: Anatomically Consistent Embeddings in Composition and Decomposition
Ziyu Zhou, Haozhe Luo, Mohammad Reza Hosseinzadeh Taher +4
Medical images acquired from standardized protocols show consistent macroscopic or microscopic anatomical structures, and these structures consist of composable/decomposable organs…