3 papers
cs.CV2025
Descrip3D: Enhancing Large Language Model-based 3D Scene Understanding with Object-Level Text Descriptions
Jintang Xue, Ganning Zhao, Jie-En Yao +5
Understanding 3D scenes goes beyond simply recognizing objects; it requires reasoning about the spatial and semantic relationships between them. Current 3D scene-language models of…
cs.CV2025
EVLF-FM: Explainable Vision Language Foundation Model for Medicine
Yang Bai, Haoran Cheng, Yang Zhou +40
Despite the promise of foundation models in medical AI, current systems remain limited - they are modality-specific and lack transparent reasoning processes, hindering clinical ado…
eess.IV2025
Multimodal, Multi-Disease Medical Imaging Foundation Model (MerMED-FM)
Yang Zhou, Chrystie Wan Ning Quek, Jun Zhou +22
Current artificial intelligence models for medical imaging are predominantly single modality and single disease. Attempts to create multimodal and multi-disease models have resulte…