4 papers
SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?
Jongmin Shin, Ka Young Kim, Eunki Cho +2
Purpose: Vision-language models (VLMs) have shown promising performance in surgical visual question answering (VQA). However, existing surgical VQA datasets often contain linguisti…
A generalizable foundation model for intraoperative understanding across surgical procedures
Kanggil Park, Yongjun Jeon, Soyoung Lim +9
In minimally invasive surgery, clinical decisions depend on real-time visual interpretation, yet intraoperative perception varies substantially across surgeons and procedures. This…
CurConMix+: A Unified Spatio-Temporal Framework for Hierarchical Surgical Workflow Understanding
Yongjun Jeon, Jongmin Shin, Kanggil Park +8
Surgical action triplet recognition aims to understand fine-grained surgical behaviors by modeling the interactions among instruments, actions, and anatomical targets. Despite its…
Towards Holistic Surgical Scene Graph
Jongmin Shin, Enki Cho, Ka Young Kim +3
Surgical scene understanding is crucial for computer-assisted intervention systems, requiring visual comprehension of surgical scenes that involves diverse elements such as surgica…