papers

Publications (109)

cs.CV2023

Why Deep Surgical Models Fail?: Revisiting Surgical Action Triplet Recognition through the Lens of Robustness

Yanqi Cheng, Lihao Liu, Shujun Wang +3

Surgical action triplet recognition provides a better understanding of the surgical scene. This task is of high relevance as it provides the surgeon with context-aware support and…

cs.SE2026

Git Context Controller: Manage the Context of LLM-based Agents like Git

Junde Wu, Minhao Hu, Jiayuan Zhu +4

Large language model (LLM) agents have demonstrated strong capabilities in long-horizon tasks by interleaving reasoning with tool use. However, as these agents scale to complex wor…

cs.AI2020

Automatic Gesture Recognition in Robot-assisted Surgery with Reinforcement Learning and Tree Search

Xiaojie Gao, Yueming Jin, Qi Dou +1

Automatic surgical gesture recognition is fundamental for improving intelligence in robot-assisted surgery, such as conducting complicated tasks of surgery surveillance and skill e…

cs.CV2025

ToolTipNet: A Segmentation-Driven Deep Learning Baseline for Surgical Instrument Tip Detection

Zijian Wu, Shuojue Yang, Yueming Jin +1

In robot-assisted laparoscopic radical prostatectomy (RALP), the location of the instrument tip is important to register the ultrasound frame with the laparoscopic camera frame. A…

cs.AI2026

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding

Linghao Meng, Qiankun Li, Junyuan Mao +7

While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding. To a…

cs.CV2020

Robust Multimodal Brain Tumor Segmentation via Feature Disentanglement and Gated Fusion

Cheng Chen, Qi Dou, Yueming Jin +3

Accurate medical image segmentation commonly requires effective learning of the complementary information from multimodal data. However, in clinical practice, we often encounter th…