Learning from a tiny dataset of manual annotations: a teacher/student approach for surgical phase recognition
arXiv:1812.00033
Abstract
Vision algorithms capable of interpreting scenes from a real-time video stream are necessary for computer-assisted surgery systems to achieve context-aware behavior. In laparoscopic procedures one particular algorithm needed for such systems is the identification of surgical phases, for which the current state of the art is a model based on a CNN-LSTM. A number of previous works using models of this kind have trained them in a fully supervised manner, requiring a fully annotated dataset. Instead, our work confronts the problem of learning surgical phase recognition in scenarios presenting scarce amounts of annotated data (under 25% of all available video recordings). We propose a teacher/student type of approach, where a strong predictor called the teacher, trained beforehand on a small dataset of ground truth-annotated videos, generates synthetic annotations for a larger dataset, which another model - the student - learns from. In our case, the teacher features a novel CNN-biLSTM-CRF architecture, designed for offline inference only. The student, on the other hand, is a CNN-LSTM capable of making real-time predictions. Results for various amounts of manually annotated videos demonstrate the superiority of the new CNN-biLSTM-CRF predictor as well as improved performance from the CNN-LSTM trained using synthetic labels generated for unannotated videos. For both offline and online surgical phase recognition with very few annotated recordings available, this new teacher/student strategy provides a valuable performance improvement by efficiently leveraging the unannotated data.
Accepted at IPCAI 2019
References in corpus (4)
- RSDNet: Learning to Predict Remaining Surgery Duration from Laparoscopic Videos Without Manual Annotations
- Less is More: Surgical Phase Recognition with Less Annotations through Self-Supervised Pre-training of CNN-LSTM Networks
- Unsupervised temporal context learning using convolutional neural networks for laparoscopic workflow analysis
- Single- and Multi-Task Architectures for Tool Presence Detection Challenge at M2CAI 2016
Cited by in corpus (13)
- Rendezvous: Attention Mechanisms for the Recognition of Surgical Action Triplets in Endoscopic Videos
- Gesture Recognition in Robotic Surgery: a Review
- CholecTriplet2021: A benchmark challenge for surgical action triplet recognition
- A Kinematic Bottleneck Approach For Pose Regression of Flexible Surgical Instruments directly from Images
- Weakly Supervised Temporal Convolutional Networks for Fine-grained Surgical Activity Recognition
- Multi-Task Temporal Convolutional Networks for Joint Recognition of Surgical Phases and Steps in Gastric Bypass Procedures
- Temporal Segmentation of Surgical Sub-tasks through Deep Learning with Multiple Data Sources
- LRTD: Long-Range Temporal Dependency based Active Learning for Surgical Workflow Recognition
- daVinciNet: Joint Prediction of Motion and Surgical State in Robot-Assisted Surgery
- Rethinking Generalization Performance of Surgical Phase Recognition with Expert-Generated Annotations
- Learning Invariant Representation of Tasks for Robust Surgical State Estimation
- Artificial Intelligence in Surgery: Neural Networks and Deep Learning
- A Novel Deep ML Architecture by Integrating Visual Simultaneous Localization and Mapping (vSLAM) into Mask R-CNN for Real-time Surgical Video Analysis