From the 1 of 7 linked papers with an AI index.
7 papers
MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models
Yitao Zhu, Mengjun Liu, Yingji Fu +2
MedARC is a training-free method that compresses redundant visual tokens in 3D medical images for vision‑language models by scoring token importance with multiple cues and merging…
IHF-Harmony: Multi-Modality Magnetic Resonance Images Harmonization using Invertible Hierarchy Flow Model
Pengli Zhu, Yitao Zhu, Haowen Pang +1
Retrospective MRI harmonization is limited by poor scalability across modalities and reliance on traveling subject datasets. To address these challenges, we introduce IHF-Harmony,…
VisionCAD: An Integration-Free Radiology Copilot Framework
Jiaming Li, Junlei Wu, Sheng Wang +7
Widespread clinical deployment of computer-aided diagnosis (CAD) systems is hindered by the challenge of integrating with existing hospital IT infrastructure. Here, we introduce Vi…
Med-LEGO: Editing and Adapting toward Generalist Medical Image Diagnosis
Yitao Zhu, Yuan Yin, Jiaming Li +5
The adoption of visual foundation models has become a common practice in computer-aided diagnosis (CAD). While these foundation models provide a viable solution for creating genera…
ReactDiff: Latent Diffusion for Facial Reaction Generation
Jiaming Li, Sheng Wang, Xin Wang +4
Given the audio-visual clip of the speaker, facial reaction generation aims to predict the listener's facial reactions. The challenge lies in capturing the relevance between video…
UniCAD: Efficient and Extendable Architecture for Multi-Task Computer-Aided Diagnosis System
Yitao Zhu, Yuan Yin, Zhenrong Shen +5
The growing complexity and scale of visual model pre-training have made developing and deploying multi-task computer-aided diagnosis (CAD) systems increasingly challenging and reso…