3 papers
cs.CV2025
ReactDiff: Latent Diffusion for Facial Reaction Generation
Jiaming Li, Sheng Wang, Xin Wang +4
Given the audio-visual clip of the speaker, facial reaction generation aims to predict the listener's facial reactions. The challenge lies in capturing the relevance between video…
cs.CV2024
MeLo: Low-rank Adaptation is Better than Fine-tuning for Medical Image Diagnosis
Yitao Zhu, Zhenrong Shen, Zihao Zhao +5
The common practice in developing computer-aided diagnosis (CAD) models based on transformer architectures usually involves fine-tuning from ImageNet pre-trained weights. However,…
eess.IV2024
Inter-slice Super-resolution of Magnetic Resonance Images by Pre-training and Self-supervised Fine-tuning
Xin Wang, Zhiyun Song, Yitao Zhu +4
In clinical practice, 2D magnetic resonance (MR) sequences are widely adopted. While individual 2D slices can be stacked to form a 3D volume, the relatively large slice spacing can…