3 papers
cs.CV2025
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
Guankun Wang, Junyi Wang, Wenjin Mo +11
Surgical scene understanding is critical for surgical training and robotic decision-making in robot-assisted surgery. Recent advances in Multimodal Large Language Models (MLLMs) ha…
cs.CV2025
EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery
Guankun Wang, Long Bai, Junyi Wang +13
Recently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted sur…
cs.CV2024
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
Tong Chen, Shuya Yang, Junyi Wang +3
Surgical video generation can enhance medical education and research, but existing methods lack fine-grained motion control and realism. We introduce SurgSora, a framework that gen…