Multitask Learning in Minimally Invasive Surgical Vision: A Review
arXiv:2401.08256 · doi:10.1016/j.media.2025.103480
Abstract
Minimally invasive surgery (MIS) has revolutionized many procedures and led to reduced recovery time and risk of patient injury. However, MIS poses additional complexity and burden on surgical teams. Data-driven surgical vision algorithms are thought to be key building blocks in the development of future MIS systems with improved autonomy. Recent advancements in machine learning and computer vision have led to successful applications in analyzing videos obtained from MIS with the promise of alleviating challenges in MIS videos. Surgical scene and action understanding encompasses multiple related tasks that, when solved individually, can be memory-intensive, inefficient, and fail to capture task relationships. Multitask learning (MTL), a learning paradigm that leverages information from multiple related tasks to improve performance and aid generalization, is well suited for fine-grained and high-level understanding of MIS data. This review provides a narrative overview of the current state-of-the-art MTL systems that leverage videos obtained from MIS. Beyond listing published approaches, we discuss the benefits and limitations of these MTL systems. Moreover, this manuscript presents an analysis of the literature for various application fields of MTL in MIS, including those with large models, highlighting notable trends, new directions of research, and developments.
Published at Medical Image Analysis
References in corpus (34)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?
- Segment Anything in Medical Images
- U-Net: Going Deeper with Nested U-Structure for Salient Object Detection
- A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Multi-Task Learning for Dense Prediction Tasks: A Survey
- Rendezvous: Attention Mechanisms for the Recognition of Surgical Action Triplets in Endoscopic Videos
- Auxiliary Tasks in Multi-task Learning
- CAI4CAI: The Rise of Contextual Artificial Intelligence in Computer Assisted Interventions
- RSDNet: Learning to Predict Remaining Surgery Duration from Laparoscopic Videos Without Manual Annotations
- 2018 Robotic Scene Segmentation Challenge
- Real-Time Instrument Segmentation in Robotic Surgery using Auxiliary Supervised Deep Adversarial Learning
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual Models
- Towards Better Surgical Instrument Segmentation in Endoscopic Vision: Multi-Angle Feature Aggregation and Contour Supervision
- Robotic Endoscope Control via Autonomous Instrument Tracking
- Stereo Dense Scene Reconstruction and Accurate Localization for Learning-Based Navigation of Laparoscope in Minimally Invasive Surgery
- CholecSeg8k: A Semantic Segmentation Dataset for Laparoscopic Cholecystectomy Based on Cholec80
- Global-Reasoned Multi-Task Learning Model for Surgical Scene Understanding
- Rendezvous in Time: An Attention-based Temporal Fusion approach for Surgical Triplet Recognition
- Learning joint segmentation of tissues and brain lesions from task-specific hetero-modal domain-shifted datasets
- Robust Medical Instrument Segmentation Challenge 2019
- The SARAS Endoscopic Surgeon Action Detection (ESAD) dataset: Challenges and methods
- Single- and Multi-Task Architectures for Surgical Workflow Challenge at M2CAI 2016
- In Defense of the Unitary Scalarization for Deep Multi-Task Learning
- LibMTL: A Python Library for Multi-Task Learning
- Do Current Multi-Task Optimization Methods in Deep Learning Even Help?
- Multitask Learning of Temporal Connectionism in Convolutional Networks using a Joint Distribution Loss Function to Simultaneously Identify Tools and Phase in Surgical Videos
- SAM Meets Robotic Surgery: An Empirical Study in Robustness Perspective
- SAR-RARP50: Segmentation of surgical instrumentation and Action Recognition on Robot-Assisted Radical Prostatectomy Challenge
- Text Promptable Surgical Instrument Segmentation with Vision-Language Models
- Revisiting Scalarization in Multi-Task Learning: A Theoretical Perspective
- Task-Agnostic Robust Representation Learning