1 paper
Chengyu Fang, Heng Guo, Zheng Jiang +3
Multimodal large language models are promising for clinical visual question answering tasks, but scaling to 3D imaging is hindered by high computational costs. Prior methods often…