7 papers
Anatomy-R1: Enhancing Anatomy Reasoning in Multimodal Large Language Models via Anatomical Similarity Curriculum and Group Diversity Augmentation
Ziyang Song, Zelin Zang, Zuyao Chen +6
Multimodal Large Language Models (MLLMs) have achieved impressive progress in natural image reasoning, yet their potential in medical imaging remains underexplored, especially in c…
NeuroABench: A Multimodal Evaluation Benchmark for Neurosurgical Anatomy Identification
Ziyang Song, Zelin Zang, Xiaofan Ye +7
Multimodal Large Language Models (MLLMs) have shown significant potential in surgical video understanding. With improved zero-shot performance and more effective human-machine inte…
Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
Lihua Zhou, Mao Ye, Shuaifeng Li +7
Vision-language models (VLMs) such as CLIP and Grounding DINO have achieved remarkable success in object recognition and detection. However, their performance often degrades under…
Text-Driven Causal Representation Learning for Source-Free Domain Generalization
Lihua Zhou, Mao Ye, Nianxin Li +7
Deep learning often struggles when training and test data distributions differ. Traditional domain generalization (DG) tackles this by including data from multiple source domains,…
Endo-4DGX: Robust Endoscopic Scene Reconstruction and Illumination Correction with Gaussian Splatting
Yiming Huang, Long Bai, Beilei Cui +7
Accurate reconstruction of soft tissue is crucial for advancing automation in image-guided robotic surgery. The recent 3D Gaussian Splatting (3DGS) techniques and their variants, 4…
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
Guankun Wang, Junyi Wang, Wenjin Mo +11
Surgical scene understanding is critical for surgical training and robotic decision-making in robot-assisted surgery. Recent advances in Multimodal Large Language Models (MLLMs) ha…