4 papers · 1 filter
JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation
Yue Xun, Junyu Liu, Qian Niu +10
We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese M…
PRA-PoE: Robust Multimodal Alzheimer's Diagnosis with Arbitrary Missing Modalities
Guangqian Yang, Ye Du, Wenlong Hou +2
Missing modalities are prevalent in real-world Alzheimer's disease (AD) assessment and pose a significant challenge to multimodal learning, particularly when the distribution of ob…
NEWTON: Agentic Planning for Physically Grounded Video Generation
Yuxiang Feng, Juncheng Wang, Chao Xu +7
Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves only 32.6% joint accuracy. We…
BrainAnytime: Anatomy-Aware Cross-Modal Pretraining for Brain Image Analysis with Arbitrary Modality Availability
Guangqian Yang, Tong Ding, Wenlong Hou +4
Clinical diagnostic workups typically follow a modality escalation pathway: after initial clinical evaluation, clinicians begin with routine structural imaging (e.g., MRI), selecti…