4 papers
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
Yao Du, Shanshan Song, Xiaomeng Li
Multimodal large language models (MLLMs) struggle with numerical regression under long-tailed target distributions. Token-level supervised fine-tuning (SFT) and point-wise regressi…
Divide-then-Diagnose: Weaving Clinician-Inspired Contexts for Ultra-Long Capsule Endoscopy Videos
Bowen Liu, Li Yang, Shanshan Song +6
Capsule endoscopy (CE) enables non-invasive gastrointestinal screening, but current CE research remains largely limited to frame-level classification and detection, leaving video-l…
DDaTR: Dynamic Difference-aware Temporal Residual Network for Longitudinal Radiology Report Generation
Shanshan Song, Hui Tang, Honglong Yang +1
Radiology Report Generation (RRG) automates the creation of radiology reports from medical imaging, enhancing the efficiency of the reporting process. Longitudinal Radiology Report…
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
Honglong Yang, Shanshan Song, Yi Qin +6
Generalist Medical AI (GMAI) systems have demonstrated expert-level performance in biomedical perception tasks, yet their clinical utility remains limited by inadequate multi-modal…