12 papers
Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs
Haiming Li, Yingsheng Liu, Jingmin Zhu +5
Multimodal LLMs apply the language model interface to visual inputs, where ordinal regression tasks such as age estimation, image quality assessment, and disease grading require au…
A Neurosymbolic Framework for Interpretable Skeleton-Based Seizure Detection via Concept-Driven Logical Reasoning
Talha Ilyas, Deval Mehta, Zongyuan Ge
Video-based seizure detection is essential for the management of epilepsy patients, offering a non-invasive complement to electroencephalography. While several deep learning approa…
EnTrust: Modeling Inter-Modal Conflict for Trustworthy Multimodal Medical Image Analysis
Dwarikanath Mahapatra, Abhijit Das, Behzad Bozorgtabar +5
Multimodal medical imaging fuses complementary anatomical and functional information, yet modalities frequently disagree in pathologically heterogeneous regions. Current segmentati…
Clinically Aligned Geometry Constraints for Robust IVUS Vessel Boundary Segmentation
Yunshu Chen, Litao Yang, Giuseppe Di Giovanni +9
Intravascular ultrasound (IVUS) lumen and external elastic membrane (EEM) segmentation is important for quantitative coronary plaque burden assessment. Errors in lumen or EEM delin…
DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making
Yize Liu, Siyuan Yan, Ming Hu +5
Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (MLLMs) facilitate interactiv…
OPGAgent: An Agent for Auditable Dental Panoramic X-ray Interpretation
Zhaolin Yu, Litao Yang, Ben Babicka +7
Orthopantomograms (OPGs) are the standard panoramic radiograph in dentistry, used for full-arch screening across multiple diagnostic tasks. While Vision Language Models (VLMs) now…