12 papers
EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining
Zhuo Deng, Ruiheng Zhang, Ziheng Zhang +23
Color fundus photography (CFP) is the mainstay of large-scale retinal screening, but its diagnostic capacity is limited by the lack of depth-resolved structure, which optical coher…
SimWeaver: Zero-Shot RGB Sim-to-Real for Deformable Manipulation
Wenkang Hu, Haoran Wang, Yitong Li +10
RGB sim-to-real for deformable manipulation has remained largely unsolved without real-world fine-tuning. We present SimWeaver, which trains zero-shot RGB VLA policies on 200 simul…
Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation
Tom Maye-Lasserre, Yitong Li, Bailiang Jian +3
Modern 3D medical vision-language models (VLMs) can generate fluent radiology-style text while exhibit critically low pathology detection and output diversity, collapsing to generi…
Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging
Yitong Li, Morteza Ghahremani, Christian Wachinger
Vision-language foundation models achieve promising performance in natural image classification, yet their direct application to medical imaging is limited by severe domain shifts,…
TimeFlow: Temporal Conditioning for Longitudinal Brain MRI Registration and Aging Analysis
Bailiang Jian, Jiazhen Pan, Yitong Li +5
Longitudinal brain analysis is essential for understanding healthy aging and identifying pathological deviations. Longitudinal registration of sequential brain MRI underpins such a…
Conditional Diffusion for 3D CT Volume Reconstruction from 2D X-rays
Martin Rath, Morteza Ghahremani, Yitong Li +3
Computed tomography (CT) provides rich 3D anatomical details but is often constrained by high radiation exposure, substantial costs, and limited availability. While standard chest…