3 papers
cs.CV2024
3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models
Hao Chen, Wei Zhao, Yingli Li +8
Medical image analysis is crucial in modern radiological diagnostics, especially given the exponential growth in medical imaging data. The demand for automated report generation sy…
cs.SD2024
A Survey of Foundation Models for Music Understanding
Wenjun Li, Ying Cai, Ziyang Wu +13
Music is essential in daily life, fulfilling emotional and entertainment needs, and connecting us personally, socially, and culturally. A better understanding of music can enhance…
cs.CV2024
Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
Chong Ma, Hanqi Jiang, Wenting Chen +10
In the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly…