71 citations · 142 across the 12 of their papers we have counts for
12 papers
MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models
Yitao Zhu, Mengjun Liu, Yingji Fu +2
Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visu…
IHF-Harmony: Multi-Modality Magnetic Resonance Images Harmonization using Invertible Hierarchy Flow Model
Pengli Zhu, Yitao Zhu, Haowen Pang +1
Retrospective MRI harmonization is limited by poor scalability across modalities and reliance on traveling subject datasets. To address these challenges, we introduce IHF-Harmony,…
VisionCAD: An Integration-Free Radiology Copilot Framework
Jiaming Li, Junlei Wu, Sheng Wang +7
Widespread clinical deployment of computer-aided diagnosis (CAD) systems is hindered by the challenge of integrating with existing hospital IT infrastructure. Here, we introduce Vi…
ReactDiff: Latent Diffusion for Facial Reaction Generation
Jiaming Li, Sheng Wang, Xin Wang +4
Given the audio-visual clip of the speaker, facial reaction generation aims to predict the listener's facial reactions. The challenge lies in capturing the relevance between video…
UniCAD: Efficient and Extendable Architecture for Multi-Task Computer-Aided Diagnosis System
Yitao Zhu, Yuan Yin, Zhenrong Shen +5
The growing complexity and scale of visual model pre-training have made developing and deploying multi-task computer-aided diagnosis (CAD) systems increasingly challenging and reso…
Med-LEGO: Editing and Adapting toward Generalist Medical Image Diagnosis
Yitao Zhu, Yuan Yin, Jiaming Li +5
The adoption of visual foundation models has become a common practice in computer-aided diagnosis (CAD). While these foundation models provide a viable solution for creating genera…