CholecTriplet2022: Show me a tool and tell me the triplet -- an endoscopic vision challenge for surgical action triplet detection
arXiv:2302.06294 · doi:10.1016/j.media.2023.102888
Abstract
Formalizing surgical activities as triplets of the used instruments, actions performed, and target anatomies is becoming a gold standard approach for surgical activity modeling. The benefit is that this formalization helps to obtain a more detailed understanding of tool-tissue interaction which can be used to develop better Artificial Intelligence assistance for image-guided surgery. Earlier efforts and the CholecTriplet challenge introduced in 2021 have put together techniques aimed at recognizing these triplets from surgical footage. Estimating also the spatial locations of the triplets would offer a more precise intraoperative context-aware decision support for computer-assisted intervention. This paper presents the CholecTriplet2022 challenge, which extends surgical action triplet modeling from recognition to detection. It includes weakly-supervised bounding box localization of every visible surgical instrument (or tool), as the key actors, and the modeling of each tool-activity in the form of <instrument, verb, target> triplet. The paper describes a baseline method and 10 new deep learning algorithms presented at the challenge to solve the task. It also provides thorough methodological comparisons of the methods, an in-depth analysis of the obtained results across multiple metrics, visual and procedural challenges; their significance, and useful insights for future research directions and applications in surgery.
MICCAI EndoVis CholecTriplet2022 challenge report. Published at Elsevier journal of Medical Image Analysis. 25 pages, 15 figures, 8 tables
References in corpus (14)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Is Space-Time Attention All You Need for Video Understanding?
- Visual Semantic Role Labeling
- Rendezvous: Attention Mechanisms for the Recognition of Surgical Action Triplets in Endoscopic Videos
- CAI4CAI: The Rise of Contextual Artificial Intelligence in Computer Assisted Interventions
- CholecTriplet2021: A benchmark challenge for surgical action triplet recognition
- CholecSeg8k: A Semantic Segmentation Dataset for Laparoscopic Cholecystectomy Based on Cholec80
- The SARAS Endoscopic Surgeon Action Detection (ESAD) dataset: Challenges and methods
- m2caiSeg: Semantic Segmentation of Laparoscopic Images using Convolutional Neural Networks
- Surgical Visual Domain Adaptation: Results from the MICCAI 2020 SurgVisDom Challenge
- Data Splits and Metrics for Method Benchmarking on Surgical Action Triplet Datasets
- Comparative Validation of Machine Learning Algorithms for Surgical Workflow and Skill Analysis with the HeiChole Benchmark
Cited by in corpus (3)
- Surgical Phase and Instrument Recognition: How to identify appropriate Dataset Splits
- Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge
- Grounding Surgical Action Triplets with Instrument Instance Segmentation: A Dataset and Target-Aware Fusion Approach