papers

Publications (42)

cs.CV2023

Preserving Privacy in Surgical Video Analysis Using Artificial Intelligence: A Deep Learning Classifier to Identify Out-of-Body Scenes in Endoscopic Videos

Joël L. Lavanchy, Armine Vardazaryan, Pietro Mascagni +3

Objective: To develop and validate a deep learning model for the identification of out-of-body images in endoscopic videos. Background: Surgical video analysis facilitates educatio…

eess.IV2022

Real-Time Artificial Intelligence Assistance for Safe Laparoscopic Cholecystectomy: Early-Stage Clinical Evaluation

Pietro Mascagni, Deepak Alapatt, Alfonso Lapergola +5

Artificial intelligence is set to be deployed in operating rooms to improve surgical care. This early-stage clinical evaluation shows the feasibility of concurrently attaining real…

cs.CV2026

DExTeR: Weakly Semi-Supervised Object Detection with Class and Instance Experts for Medical Imaging

Adrien Meyer, Didier Mutter, Nicolas Padoy

Detecting anatomical landmarks in medical imaging is essential for diagnosis and intervention guidance. However, object detection models rely on costly bounding box annotations, li…

cs.CV2024

Overcoming Dimensional Collapse in Self-supervised Contrastive Learning for Medical Image Segmentation

Jamshid Hassanpour, Vinkle Srivastav, Didier Mutter +1

Self-supervised learning (SSL) approaches have achieved great success when the amount of labeled data is limited. Within SSL, models learn robust feature representations by solving…

cs.CV2016

Single- and Multi-Task Architectures for Tool Presence Detection Challenge at M2CAI 2016

Andru P. Twinanda, Didier Mutter, Jacques Marescaux +2

The tool presence detection challenge at M2CAI 2016 consists of identifying the presence/absence of seven surgical tools in the images of cholecystectomy videos. Here, we propose t…

cs.CV2016

EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos

Andru P. Twinanda, Sherif Shehata, Didier Mutter +3

Surgical workflow recognition has numerous potential medical applications, such as the automatic indexing of surgical video databases and the optimization of real-time operating ro…

cs.CV2026

Self-Supervised Uncalibrated Multi-View Video Anonymization in the Operating Room

Keqi Chen, Vinkle Srivastav, Armine Vardazaryan +3

Privacy preservation is a prerequisite for using video data in Operating Room (OR) research. Effective anonymization relies on the exhaustive localization of every individual; even…

cs.CV2019

Weakly Supervised Convolutional LSTM Approach for Tool Tracking in Laparoscopic Videos

Chinedu Innocent Nwoye, Didier Mutter, Jacques Marescaux +1

Purpose: Real-time surgical tool tracking is a core component of the future intelligent operating room (OR), because it is highly instrumental to analyze and understand the surgica…

cs.CV2023

Weakly Supervised Temporal Convolutional Networks for Fine-grained Surgical Activity Recognition

Sanat Ramesh, Diego Dall'Alba, Cristians Gonzalez +6

Automatic recognition of fine-grained surgical activities, called steps, is a challenging but crucial task for intelligent intra-operative computer assistance. The development of c…

cs.CV2018

Weakly-Supervised Learning for Tool Localization in Laparoscopic Videos

Armine Vardazaryan, Didier Mutter, Jacques Marescaux +1

Surgical tool localization is an essential task for the automatic analysis of endoscopic videos. In the literature, existing methods for tool localization, tracking and segmentatio…

cs.CV2026

Where are they looking in the operating room?

Keqi Chen, Séraphin Baributsa, Lilien Schewski +5

Purpose: Gaze-following, the task of inferring where individuals are looking, has been widely studied in computer vision, advancing research in visual attention modeling, social sc…

cs.CV2024

The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark

Aditya Murali, Deepak Alapatt, Pietro Mascagni +8

This technical report provides a detailed overview of Endoscapes, a dataset of laparoscopic cholecystectomy (LC) videos with highly intricate annotations targeted at automated asse…

cs.CV2026

S4M: 4-points to Segment Anything

Adrien Meyer, Lorenzo Arboit, Giuseppe Massimiani +3

Purpose: The Segment Anything Model (SAM) promises to ease the annotation bottleneck in medical segmentation, but overlapping anatomy and blurred boundaries make its point prompts…

eess.IV2020

Recognition of Instrument-Tissue Interactions in Endoscopic Videos via Action Triplets

Chinedu Innocent Nwoye, Cristians Gonzalez, Tong Yu +4

Recognition of surgical activity is an essential component to develop context-aware decision support for the operating room. In this work, we tackle the recognition of fine-grained…

cs.CV2022

CholecTriplet2021: A benchmark challenge for surgical action triplet recognition

Chinedu Innocent Nwoye, Deepak Alapatt, Tong Yu +59

Context-aware decision support in the operating room can foster surgical safety and efficiency by leveraging real-time feedback from surgical workflow analysis. Most existing works…

cs.CV2023

Encoding Surgical Videos as Latent Spatiotemporal Graphs for Object and Anatomy-Driven Reasoning

Aditya Murali, Deepak Alapatt, Pietro Mascagni +5

Recently, spatiotemporal graphs have emerged as a concise and elegant manner of representing video clips in an object-centric fashion, and have shown to be useful for downstream ta…

cs.CV2018

RSDNet: Learning to Predict Remaining Surgery Duration from Laparoscopic Videos Without Manual Annotations

Andru Putra Twinanda, Gaurav Yengera, Didier Mutter +2

Accurate surgery duration estimation is necessary for optimal OR planning, which plays an important role in patient comfort and safety as well as resource optimization. It is, howe…

cs.CV2016

Single- and Multi-Task Architectures for Surgical Workflow Challenge at M2CAI 2016

Andru P. Twinanda, Didier Mutter, Jacques Marescaux +2

The surgical workflow challenge at M2CAI 2016 consists of identifying 8 surgical phases in cholecystectomy procedures. Here, we propose to use deep architectures that are based on…

cs.CV2018

Less is More: Surgical Phase Recognition with Less Annotations through Self-Supervised Pre-training of CNN-LSTM Networks

Gaurav Yengera, Didier Mutter, Jacques Marescaux +1

Real-time algorithms for automatically recognizing surgical phases are needed to develop systems that can provide assistance to surgeons, enable better management of operating room…

cs.CV2025

Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition

Idris Hamoud, Vinkle Srivastav, Muhammad Abdullah Jamal +3

Understanding the workflow of surgical procedures in complex operating rooms requires a deep understanding of the interactions between clinicians and their environment. Surgical ac…

q-bio.TO2021

Intraoperative time out to promote the implementation of the critical view of safety in laparoscopic cholecystectomy: a video-based assessment of 343 procedures

Pietro Mascagni, Maria Rita Rodriguez-Luna, Takeshi Urade +7

Background: The critical view of safety (CVS) is poorly adopted in surgical practices although it is ubiquitously recommended to prevent major bile duct injuries during laparoscopi…

cs.LG2020

Learning from a tiny dataset of manual annotations: a teacher/student approach for surgical phase recognition

Tong Yu, Didier Mutter, Jacques Marescaux +1

Vision algorithms capable of interpreting scenes from a real-time video stream are necessary for computer-assisted surgery systems to achieve context-aware behavior. In laparoscopi…

cs.CV2025

State-Change Learning for Prediction of Future Events in Endoscopic Videos

Saurav Sharma, Chinedu Innocent Nwoye, Didier Mutter +1

Surgical future prediction, driven by real-time AI analysis of surgical video, is critical for operating room safety and efficiency. It provides actionable insights into upcoming e…

eess.IV2023

TRUSTED: The Paired 3D Transabdominal Ultrasound and CT Human Data for Kidney Segmentation and Registration Research

William Ndzimbong, Cyril Fourniol, Loic Themyr +10

Inter-modal image registration (IMIR) and image segmentation with abdominal Ultrasound (US) data has many important clinical applications, including image-guided surgery, automatic…

cs.CV2025

CycleSAM: Few-Shot Surgical Scene Segmentation with Cycle- and Scene-Consistent Feature Matching

Aditya Murali, Farahdiba Zarin, Adrien Meyer +3

Surgical image segmentation is highly challenging, primarily due to scarcity of annotated data. Generalist prompted segmentation models like the Segment-Anything Model (SAM) can he…

cs.CV2025

fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models

Saurav Sharma, Didier Mutter, Nicolas Padoy

While vision-language models like CLIP have advanced zero-shot surgical phase recognition, they struggle with fine-grained surgical activities, especially action triplets. This lim…

eess.IV2023

Live Laparoscopic Video Retrieval with Compressed Uncertainty

Tong Yu, Pietro Mascagni, Juan Verde +3

Searching through large volumes of medical data to retrieve relevant information is a challenging yet crucial task for clinical care. However the primitive and most common approach…

cs.CV2021

Multi-Task Temporal Convolutional Networks for Joint Recognition of Surgical Phases and Steps in Gastric Bypass Procedures

Sanat Ramesh, Diego Dall'Alba, Cristians Gonzalez +6

Purpose: Automatic segmentation and classification of surgical activity is crucial for providing advanced support in computer-assisted interventions and autonomous functionalities…

cs.CV2026

SPIRIT: Spatio-temporal Pairwise Relational Modeling of Instrument-Tissue Interactions for Surgical Action Triplet Recognition

Saurav Sharma, Lorenzo Arboit, Nabani Banik +11

Fine-grained understanding of surgical activity is essential for context-aware assistance in the operating room, including safety monitoring, adverse event identification, and skil…

cs.CV2023

ST(OR)2: Spatio-Temporal Object Level Reasoning for Activity Recognition in the Operating Room

Idris Hamoud, Muhammad Abdullah Jamal, Vinkle Srivastav +3

Surgical robotics holds much promise for improving patient safety and clinician experience in the Operating Room (OR). However, it also comes with new challenges, requiring strong…

cs.CV2023

Surgical Action Triplet Detection by Mixed Supervised Learning of Instrument-Tissue Interactions

Saurav Sharma, Chinedu Innocent Nwoye, Didier Mutter +1

Surgical action triplets describe instrument-tissue interactions as (instrument, verb, target) combinations, thereby supporting a detailed analysis of surgical scene activities and…

cs.CV2022

Rendezvous: Attention Mechanisms for the Recognition of Surgical Action Triplets in Endoscopic Videos

Chinedu Innocent Nwoye, Tong Yu, Cristians Gonzalez +5

Out of all existing frameworks for surgical workflow analysis in endoscopic videos, action triplet recognition stands out as the only one aiming to provide truly fine-grained and c…

cs.CV2025

When do they StOP?: A First Step Towards Automatically Identifying Team Communication in the Operating Room

Keqi Chen, Lilien Schewski, Vinkle Srivastav +5

Purpose: Surgical performance depends not only on surgeons' technical skills but also on team communication within and across the different professional groups present during the o…

cs.CV2019

Future-State Predicting LSTM for Early Surgery Type Recognition

Siddharth Kannan, Gaurav Yengera, Didier Mutter +2

This work presents a novel approach for the early recognition of the type of a laparoscopic surgery from its video. Early recognition algorithms can be beneficial to the developmen…

cs.CV2025

Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes

Keqi Chen, Vinkle Srivastav, Didier Mutter +1

Multi-view person association is a fundamental step towards multi-view analysis of human activities. Although the person re-identification features have been proven effective, they…

cs.CV2021

Temporally Constrained Neural Networks (TCNN): A framework for semi-supervised video semantic segmentation

Deepak Alapatt, Pietro Mascagni, Armine Vardazaryan +7

A major obstacle to building models for effective semantic segmentation, and particularly video semantic segmentation, is a lack of large and well annotated datasets. This bottlene…

cs.CV2025

Early Operative Difficulty Assessment in Laparoscopic Cholecystectomy via Snapshot-Centric Video Analysis

Saurav Sharma, Maria Vannucci, Leonardo Pestana Legori +6

Purpose: Laparoscopic cholecystectomy (LC) operative difficulty (LCOD) is highly variable and influences outcomes. Despite extensive LC studies in surgical workflow analysis, limit…

cs.CV2023

Challenges in Multi-centric Generalization: Phase and Step Recognition in Roux-en-Y Gastric Bypass Surgery

Joel L. Lavanchy, Sanat Ramesh, Diego Dall'Alba +7

Most studies on surgical activity recognition utilizing Artificial intelligence (AI) have focused mainly on recognizing one type of activity from small and mono-centric surgical vi…

eess.IV2023

CholecTriplet2022: Show me a tool and tell me the triplet -- an endoscopic vision challenge for surgical action triplet detection

Chinedu Innocent Nwoye, Tong Yu, Saurav Sharma +46

Formalizing surgical activities as triplets of the used instruments, actions performed, and target anatomies is becoming a gold standard approach for surgical activity modeling. Th…

cs.CV2023

Rendezvous in Time: An Attention-based Temporal Fusion approach for Surgical Triplet Recognition

Saurav Sharma, Chinedu Innocent Nwoye, Didier Mutter +1

One of the recent advances in surgical AI is the recognition of surgical activities as triplets of (instrument, verb, target). Albeit providing detailed information for computer-as…

cs.CV2023

Latent Graph Representations for Critical View of Safety Assessment

Aditya Murali, Deepak Alapatt, Pietro Mascagni +5

Assessing the critical view of safety in laparoscopic cholecystectomy requires accurate identification and localization of key anatomical structures, reasoning about their geometri…

eess.IV2025

UltraSam: A Foundation Model for Ultrasound using Large Open-Access Segmentation Datasets

Adrien Meyer, Aditya Murali, Farahdiba Zarin +2

Purpose: Automated ultrasound image analysis is challenging due to anatomical complexity and limited annotated data. To tackle this, we take a data-centric approach, assembling the…