Rendezvous: Attention Mechanisms for the Recognition of Surgical Action Triplets in Endoscopic Videos
arXiv:2109.03223 · doi:10.1016/j.media.2022.102433
Abstract
Out of all existing frameworks for surgical workflow analysis in endoscopic videos, action triplet recognition stands out as the only one aiming to provide truly fine-grained and comprehensive information on surgical activities. This information, presented as <instrument, verb, target> combinations, is highly challenging to be accurately identified. Triplet components can be difficult to recognize individually; in this task, it requires not only performing recognition simultaneously for all three triplet components, but also correctly establishing the data association between them. To achieve this task, we introduce our new model, the Rendezvous (RDV), which recognizes triplets directly from surgical videos by leveraging attention at two different levels. We first introduce a new form of spatial attention to capture individual action triplet components in a scene; called Class Activation Guided Attention Mechanism (CAGAM). This technique focuses on the recognition of verbs and targets using activations resulting from instruments. To solve the association problem, our RDV model adds a new form of semantic attention inspired by Transformer networks; called Multi-Head of Mixed Attention (MHMA). This technique uses several cross and self attentions to effectively capture relationships between instruments, verbs, and targets. We also introduce CholecT50 - a dataset of 50 endoscopic videos in which every frame has been annotated with labels from 100 triplet classes. Our proposed RDV model significantly improves the triplet prediction mean AP by over 9% compared to the state-of-the-art methods on this dataset.
21 pages, 11 figures, 19 tables, 1 video. Accepted at Elsevier Journal of Medical Image Analysis. Supplementary video available at: https://youtu.be/d_yHdJtCa98
References in corpus (12)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- Is Space-Time Attention All You Need for Video Understanding?
- CAI4CAI: The Rise of Contextual Artificial Intelligence in Computer Assisted Interventions
- CPTR: Full Transformer Network for Image Captioning
- OperA: Attention-Regularized Transformers for Surgical Phase Recognition
- Temporal Attention Model for Neural Machine Translation
- The SARAS Endoscopic Surgeon Action Detection (ESAD) dataset: Challenges and methods
- HOTR: End-to-End Human-Object Interaction Detection with Transformers
- Trans-SVNet: Accurate Phase Recognition from Surgical Videos via Hybrid Embedding Aggregation Transformer
- Multi-Task Temporal Convolutional Networks for Joint Recognition of Surgical Phases and Steps in Gastric Bypass Procedures
- End-to-End Attention-based Image Captioning
Cited by in corpus (15)
- CholecTriplet2021: A benchmark challenge for surgical action triplet recognition
- Towards Holistic Surgical Scene Understanding
- CholecTriplet2022: Show me a tool and tell me the triplet -- an endoscopic vision challenge for surgical action triplet detection
- Rendezvous in Time: An Attention-based Temporal Fusion approach for Surgical Triplet Recognition
- CholecInstanceSeg: A Tool Instance Segmentation Dataset for Laparoscopic Surgery
- Data Splits and Metrics for Method Benchmarking on Surgical Action Triplet Datasets
- Evaluating the Task Generalization of Temporal Convolutional Networks for Surgical Gesture and Motion Recognition using Kinematic Data
- Surgical Text-to-Image Generation
- Multitask Learning in Minimally Invasive Surgical Vision: A Review
- SegMatch: A semi-supervised learning method for surgical instrument segmentation
- COMPASS: A Formal Framework and Aggregate Dataset for Generalized Surgical Procedure Modeling
- Adaptive transfer learning for surgical tool presence detection in laparoscopic videos through gradual freezing fine-tuning
- Learning dissection trajectories from expert surgical videos via imitation learning with equivariant diffusion
- PoCaPNet: A Novel Approach for Surgical Phase Recognition Using Speech and X-Ray Images
- Grounding Surgical Action Triplets with Instrument Instance Segmentation: A Dataset and Target-Aware Fusion Approach