papers

Publications (63)

cs.CV2024

Surgical-DeSAM: Decoupling SAM for Instrument Segmentation in Robotic Surgery

Yuyang Sheng, Sophia Bano, Matthew J. Clarkson +1

Purpose: The recent Segment Anything Model (SAM) has demonstrated impressive performance with point, text or bounding box prompts, in various applications. However, in safety-criti…

eess.IV2022

2020 CATARACTS Semantic Segmentation Challenge

Imanol Luengo, Maria Grammatikopoulou, Rahim Mohammadi +37

Surgical scene segmentation is essential for anatomy and instrument localization which can be further used to assess tissue-instrument interactions during a surgical procedure. In…

cs.CV2026

Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025

Aneeq Zia, Max Berniker, Rogerio Garcia Nespolo +153

Robotic assisted (RA) surgery promises to transform surgical intervention. Intuitive Surgical is committed to fostering these changes and the machine learning models and algorithms…

eess.IV2024

EndoDAC: Efficient Adapting Foundation Model for Self-Supervised Depth Estimation from Any Endoscopic Camera

Beilei Cui, Mobarakol Islam, Long Bai +2

Depth estimation plays a crucial role in various tasks within endoscopic surgery, including navigation, surface reconstruction, and augmented reality visualization. Despite the sig…

cs.CV2024

Privacy-Preserving Synthetic Continual Semantic Segmentation for Robotic Surgery

Mengya Xu, Mobarakol Islam, Long Bai +1

Deep Neural Networks (DNNs) based semantic segmentation of the robotic instruments and tissues can enhance the precision of surgical activities in robot-assisted surgery. However,…

cs.CV2021

Class-Incremental Domain Adaptation with Smoothing and Calibration for Surgical Report Generation

Mengya Xu, Mobarakol Islam, Chwee Ming Lim +1

Generating surgical reports aimed at surgical scene understanding in robot-assisted surgery can contribute to documenting entry tasks and post-operative analysis. Despite the impre…

cs.CV2021

Spatially Varying Label Smoothing: Capturing Uncertainty from Expert Annotations

Mobarakol Islam, Ben Glocker

The task of image segmentation is inherently noisy due to ambiguities regarding the exact location of boundaries between anatomical structures. We argue that this information can b…

eess.IV2024

Illumination Histogram Consistency Metric for Quantitative Assessment of Video Sequences

Long Chen, Mobarakol Islam, Matt Clarkson +1

The advances in deep generative models have greatly accelerate the process of video procession such as video enhancement and synthesis. Learning spatio-temporal video models requir…

cs.CV2021

ST-MTL: Spatio-Temporal Multitask Learning Model to Predict Scanpath While Tracking Instruments in Robotic Surgery

Mobarakol Islam, Vibashan VS, Chwee Ming Lim +1

Representation learning of the task-oriented attention while tracking instrument holds vast potential in image-guided robotic surgery. Incorporating cognitive ability to automate t…

cs.CV2019

Learning Where to Look While Tracking Instruments in Robot-assisted Surgery

Mobarakol Islam, Yueyuan Li, Hongliang Ren

Directing of the task-specific attention while tracking instrument in surgery holds great potential in robot-assisted intervention. For this purpose, we propose an end-to-end train…

cs.CV2024

DARES: Depth Anything in Robotic Endoscopic Surgery with Self-supervised Vector-LoRA of the Foundation Model

Mona Sheikh Zeinoddin, Chiara Lena, Jiongqi Qu +11

Robotic-assisted surgery (RAS) relies on accurate depth estimation for 3D reconstruction and visualization. While foundation models like Depth Anything Models (DAM) show promise, d…

eess.IV2023

Robustness Stress Testing in Medical Image Classification

Mobarakol Islam, Zeju Li, Ben Glocker

Deep neural networks have shown impressive performance for image-based disease detection. Performance is commonly evaluated through clinical validation on independent test sets to…

eess.IV2024

EndoUIC: Promptable Diffusion Transformer for Unified Illumination Correction in Capsule Endoscopy

Long Bai, Tong Chen, Qiaozhi Tan +10

Wireless Capsule Endoscopy (WCE) is highly valued for its non-invasive and painless approach, though its effectiveness is compromised by uneven illumination from hardware constrain…

cs.CV2023

SurgicalGPT: End-to-End Language-Vision GPT for Visual Question Answering in Surgery

Lalithkumar Seenivasan, Mobarakol Islam, Gokul Kannan +1

Advances in GPT-based large language models (LLMs) are revolutionizing natural language processing, exponentially increasing its use across various domains. Incorporating uni-direc…

cs.CV2024

Surgical-VQLA++: Adversarial Contrastive Learning for Calibrated Robust Visual Question-Localized Answering in Robotic Surgery

Long Bai, Guankun Wang, Mobarakol Islam +3

Medical visual question answering (VQA) bridges the gap between visual information and clinical decision-making, enabling doctors to extract understanding from clinical images and…

cs.CV2023

Surgical-VQLA: Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery

Long Bai, Mobarakol Islam, Lalithkumar Seenivasan +1

Despite the availability of computer-aided simulators and recorded videos of surgical procedures, junior residents still heavily rely on experts to answer their queries. However, e…

eess.IV2021

Glioma Prognosis: Segmentation of the Tumor and Survival Prediction using Shape, Geometric and Clinical Information

Mobarakol Islam, V Jeya Maria Jose, Hongliang Ren

Segmentation of brain tumor from magnetic resonance imaging (MRI) is a vital process to improve diagnosis, treatment planning and to study the difference between subjects with tumo…

cs.CV2022

Confidence-Aware Paced-Curriculum Learning by Label Smoothing for Surgical Scene Understanding

Mengya Xu, Mobarakol Islam, Ben Glocker +1

Curriculum learning and self-paced learning are the training strategies that gradually feed the samples from easy to more complex. They have captivated increasing attention due to…

cs.CV2020

Real-Time Instrument Segmentation in Robotic Surgery using Auxiliary Supervised Deep Adversarial Learning

Mobarakol Islam, Daniel A. Atputharuban, Ravikiran Ramesh +1

Robot-assisted surgery is an emerging technology which has undergone rapid growth with the development of robotics and imaging systems. Innovations in vision, haptics and accurate…

cs.CV2025

Learning to Efficiently Adapt Foundation Models for Self-Supervised Endoscopic 3D Scene Reconstruction from Any Cameras

Beilei Cui, Long Bai, Mobarakol Islam +8

Accurate 3D scene reconstruction is essential for numerous medical tasks. Given the challenges in obtaining ground truth data, there has been an increasing focus on self-supervised…

cs.CV2024

EndoOOD: Uncertainty-aware Out-of-distribution Detection in Capsule Endoscopy Diagnosis

Qiaozhi Tan, Long Bai, Guankun Wang +2

Wireless capsule endoscopy (WCE) is a non-invasive diagnostic procedure that enables visualization of the gastrointestinal (GI) tract. Deep learning-based methods have shown effect…

cs.CV2021

Class-Distribution-Aware Calibration for Long-Tailed Visual Recognition

Mobarakol Islam, Lalithkumar Seenivasan, Hongliang Ren +1

Despite impressive accuracy, deep neural networks are often miscalibrated and tend to overly confident predictions. Recent techniques like temperature scaling (TS) and label smooth…

eess.IV2021

Glioblastoma Multiforme Prognosis: MRI Missing Modality Generation, Segmentation and Radiogenomic Survival Prediction

Mobarakol Islam, Navodini Wijethilake, Hongliang Ren

The accurate prognosis of Glioblastoma Multiforme (GBM) plays an essential role in planning correlated surgeries and treatments. The conventional models of survival prediction rely…

cs.CV2020

AP-MTL: Attention Pruned Multi-task Learning Model for Real-time Instrument Detection and Segmentation in Robot-assisted Surgery

Mobarakol Islam, Vibashan VS, Hongliang Ren

Surgical scene understanding and multi-tasking learning are crucial for image-guided robotic surgery. Training a real-time robotic system for the detection and segmentation of high…

cs.IR2024

LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic Surgery

Yuyang Du, Kexin Chen, Yue Zhan +7

Visual question answering (VQA) is crucial for promoting surgical education. In practice, the needs of trainees are constantly evolving, such as learning more surgical types, adapt…

cs.CV2023

CAT-ViL: Co-Attention Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery

Long Bai, Mobarakol Islam, Hongliang Ren

Medical students and junior surgeons often rely on senior surgeons and specialists to answer their questions when learning surgery. However, experts are often busy with clinical an…

cs.CV2023

Revisiting Distillation for Continual Learning on Visual Question Localized-Answering in Robotic Surgery

Long Bai, Mobarakol Islam, Hongliang Ren

The visual-question localized-answering (VQLA) system can serve as a knowledgeable assistant in surgical education. Except for providing text-based answers, the VQLA system can hig…

cs.CV2025

PitVQA++: Vector Matrix-Low-Rank Adaptation for Open-Ended Visual Question Answering in Pituitary Surgery

Runlong He, Danyal Z. Khan, Evangelos B. Mazomenos +4

Vision-Language Models (VLMs) in visual question answering (VQA) offer a unique opportunity to enhance intra-operative decision-making, promote intuitive interactions, and signific…

q-bio.QM2020

Radiogenomics of Glioblastoma: Identification of Radiomics associated with Molecular Subtypes

Navodini Wijethilake, Mobarakol Islam, Dulani Meedeniya +3

Glioblastoma is the most malignant type of central nervous system tumor with GBM subtypes cleaved based on molecular level gene alterations. These alterations are also happened to…

cs.CV2024

SimCol3D -- 3D Reconstruction during Colonoscopy Challenge

Anita Rau, Sophia Bano, Yueming Jin +19

Colorectal cancer is one of the most common cancers in the world. While colonoscopy is an effective screening technique, navigating an endoscope through the colon to detect polyps…

cs.CV2022

Rethinking Surgical Captioning: End-to-End Window-Based MLP Transformer Using Patches

Mengya Xu, Mobarakol Islam, Hongliang Ren

Surgical captioning plays an important role in surgical instruction prediction and report generation. However, the majority of captioning models still rely on the heavy computation…

cs.CV2019

Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation, Progression Assessment, and Overall Survival Prediction in the BRATS Challenge

Spyridon Bakas, Mauricio Reyes, Andras Jakab +421

Gliomas are the most common primary brain malignancies, with different degrees of aggressiveness, variable prognosis and various heterogeneous histologic sub-regions, i.e., peritum…

cs.CV2024

Endo-4DGS: Endoscopic Monocular Scene Reconstruction with 4D Gaussian Splatting

Yiming Huang, Beilei Cui, Long Bai +4

In the realm of robot-assisted minimally invasive surgery, dynamic scene reconstruction can significantly enhance downstream tasks and improve surgical outcomes. Neural Radiance Fi…

cs.CV2020

Learning and Reasoning with the Graph Structure Representation in Robotic Surgery

Mobarakol Islam, Lalithkumar Seenivasan, Lim Chwee Ming +1

Learning to infer graph representations and performing spatial reasoning in a complex surgical environment can play a vital role in surgical scene understanding in robotic surgery.…

cs.CV2024

SAM 2 in Robotic Surgery: An Empirical Evaluation for Robustness and Generalization in Surgical Video Segmentation

Jieming Yu, An Wang, Wenzhen Dong +5

The recent Segment Anything Model (SAM) 2 has demonstrated remarkable foundational competence in semantic segmentation, with its memory mechanism and mask decoder further addressin…

eess.IV2023

SAM Meets Robotic Surgery: An Empirical Study on Generalization, Robustness and Adaptation

An Wang, Mobarakol Islam, Mengya Xu +2

The Segment Anything Model (SAM) serves as a fundamental model for semantic segmentation and demonstrates remarkable generalization capabilities across a wide range of downstream s…

cs.CV2023

Paced-Curriculum Distillation with Prediction and Label Uncertainty for Image Segmentation

Mobarakol Islam, Lalithkumar Seenivasan, S. P. Sharan +4

Purpose: In curriculum learning, the idea is to train on easier samples first and gradually increase the difficulty, while in self-paced learning, a pacing function defines the spe…

cs.CV2022

Surgical-VQA: Visual Question Answering in Surgical Scenes using Transformer

Lalithkumar Seenivasan, Mobarakol Islam, Adithya K Krishna +1

Visual question answering (VQA) in surgery is largely unexplored. Expert surgeons are scarce and are often overloaded with clinical and academic workloads. This overload often limi…

eess.IV2023

SME: Spatial-Spectral Mutual Teaching and Ensemble Learning for Scribble-supervised Polyp Segmentation

An Wang, Mengya Xu, Yang Zhang +2

Fully-supervised polyp segmentation has accomplished significant triumphs over the years in advancing the early diagnosis of colorectal cancer. However, label-efficient solutions f…

cs.CV2025

Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement

Long Bai, Boyi Ma, Ruohan Wang +8

Surgical workflow recognition is vital for automating tasks, supporting decision-making, and training novice surgeons, ultimately improving patient safety and standardizing procedu…

eess.IV2023

LLCaps: Learning to Illuminate Low-Light Capsule Endoscopy with Curved Wavelet Attention and Reverse Diffusion

Long Bai, Tong Chen, Yanan Wu +3

Wireless capsule endoscopy (WCE) is a painless and non-invasive diagnostic tool for gastrointestinal (GI) diseases. However, due to GI anatomical constraints and hardware manufactu…

cs.CV2022

Angular Gap: Reducing the Uncertainty of Image Difficulty through Model Calibration

Bohua Peng, Mobarakol Islam, Mei Tu

Curriculum learning needs example difficulty to proceed from easy to hard. However, the credibility of image difficulty is rarely investigated, which can seriously affect the effec…

cs.CV2022

CholecTriplet2021: A benchmark challenge for surgical action triplet recognition

Chinedu Innocent Nwoye, Deepak Alapatt, Tong Yu +59

Context-aware decision support in the operating room can foster surgical safety and efficiency by leveraging real-time feedback from surgical workflow analysis. Most existing works…

eess.IV2023

SAM Meets Robotic Surgery: An Empirical Study in Robustness Perspective

An Wang, Mobarakol Islam, Mengya Xu +2

Segment Anything Model (SAM) is a foundation model for semantic segmentation and shows excellent generalization capability with the prompts. In this empirical study, we investigate…

cs.CV2024

Surgical-DINO: Adapter Learning of Foundation Models for Depth Estimation in Endoscopic Surgery

Beilei Cui, Mobarakol Islam, Long Bai +1

Purpose: Depth estimation in robotic surgery is vital in 3D reconstruction, surgical navigation and augmented reality visualization. Although the foundation model exhibits outstand…

eess.IV2022

Class Balanced PixelNet for Neurological Image Segmentation

Mobarakol Islam, Hongliang Ren

In this paper, we propose an automatic brain tumor segmentation approach (e.g., PixelNet) using a pixel-level convolutional neural network (CNN). The model extracts feature from mu…

eess.IV2022

Ischemic Stroke Lesion Segmentation Using Adversarial Learning

Mobarakol Islam, N Rajiv Vaidyanathan, V Jeya Maria Jose +1

Ischemic stroke occurs through a blockage of clogged blood vessels supplying blood to the brain. Segmentation of the stroke lesion is vital to improve diagnosis, outcome assessment…

cs.CV2025

Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery

Guankun Wang, Long Bai, Wan Jun Nah +7

Recent advancements in Surgical Visual Question Answering (Surgical-VQA) and related region grounding have shown great promise for robotic and medical applications, addressing the…

eess.IV2024

Transferring Knowledge from High-Quality to Low-Quality MRI for Adult Glioma Diagnosis

Yanguang Zhao, Long Bai, Zhaoxi Zhang +3

Glioma, a common and deadly brain tumor, requires early diagnosis for improved prognosis. However, low-quality Magnetic Resonance Imaging (MRI) technology in Sub-Saharan Africa (SS…

cs.CL2022

Bi-Link: Bridging Inductive Link Predictions from Text via Contrastive Learning of Transformers and Prompts

Bohua Peng, Shihao Liang, Mobarakol Islam

Inductive knowledge graph completion requires models to comprehend the underlying semantics and logic patterns of relations. With the advance of pretrained language models, recent…

eess.IV2023

Curriculum-Based Augmented Fourier Domain Adaptation for Robust Medical Image Segmentation

An Wang, Mobarakol Islam, Mengya Xu +1

Accurate and robust medical image segmentation is fundamental and crucial for enhancing the autonomy of computer-aided diagnosis and intervention systems. Medical data collection n…

eess.IV2021

Brain Tumor Segmentation and Survival Prediction using 3D Attention UNet

Mobarakol Islam, Vibashan VS, V Jeya Maria Jose +3

In this work, we develop an attention convolutional neural network (CNN) to segment brain tumors from Magnetic Resonance Images (MRI). Further, we predict the survival rate using v…

eess.IV2023

Generalizing Surgical Instruments Segmentation to Unseen Domains with One-to-Many Synthesis

An Wang, Mobarakol Islam, Mengya Xu +1

Despite their impressive performance in various surgical scene understanding tasks, deep learning-based methods are frequently hindered from deploying to real-world surgical applic…

cs.CV2024

PitVQA: Image-grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery

Runlong He, Mengya Xu, Adrito Das +6

Visual Question Answering (VQA) within the surgical domain, utilizing Large Language Models (LLMs), offers a distinct opportunity to improve intra-operative decision-making and fac…

cs.CV2022

Frequency Dropout: Feature-Level Regularization via Randomized Filtering

Mobarakol Islam, Ben Glocker

Deep convolutional neural networks have shown remarkable performance on various computer vision tasks, and yet, they are susceptible to picking up spurious correlations from the tr…

cs.CV2024

OSSAR: Towards Open-Set Surgical Activity Recognition in Robot-assisted Surgery

Long Bai, Guankun Wang, Jie Wang +6

In the realm of automated robotic surgery and computer-assisted interventions, understanding robotic surgical activities stands paramount. Existing algorithms dedicated to surgical…

cs.CV2022

Estimating Model Performance under Domain Shifts with Class-Specific Confidence Scores

Zeju Li, Konstantinos Kamnitsas, Mobarakol Islam +2

Machine learning models are typically deployed in a test setting that differs from the training setting, potentially leading to decreased model performance because of domain shift.…

eess.IV2022

Global-Reasoned Multi-Task Learning Model for Surgical Scene Understanding

Lalithkumar Seenivasan, Sai Mitheran, Mobarakol Islam +1

Global and local relational reasoning enable scene understanding models to perform human-like scene analysis and understanding. Scene understanding enables better semantic segmenta…

cs.CV2022

Rethinking Surgical Instrument Segmentation: A Background Image Can Be All You Need

An Wang, Mobarakol Islam, Mengya Xu +1

Data diversity and volume are crucial to the success of training deep learning models, while in the medical imaging field, the difficulty and cost of data collection and annotation…

eess.IV2023

Landmark Detection using Transformer Toward Robot-assisted Nasal Airway Intubation

Tianhang Liu, Hechen Li, Long Bai +4

Robot-assisted airway intubation application needs high accuracy in locating targets and organs. Two vital landmarks, nostrils and glottis, can be detected during the intubation to…

cs.CV2024

SurgicalGS: Dynamic 3D Gaussian Splatting for Accurate Robotic-Assisted Surgical Scene Reconstruction

Jialei Chen, Xin Zhang, Mobarakol Islam +4

Accurate 3D reconstruction of dynamic surgical scenes from endoscopic video is essential for robotic-assisted surgery. While recent 3D Gaussian Splatting methods have shown promise…

cs.AI2022

Task-Aware Asynchronous Multi-Task Model with Class Incremental Contrastive Learning for Surgical Scene Understanding

Lalithkumar Seenivasan, Mobarakol Islam, Mengya Xu +2

Purpose: Surgery scene understanding with tool-tissue interaction recognition and automatic report generation can play an important role in intra-operative guidance, decision-makin…

cs.RO2021

Learning Domain Adaptation with Model Calibration for Surgical Report Generation in Robotic Surgery

Mengya Xu, Mobarakol Islam, Chwee Ming Lim +1

Generating a surgical report in robot-assisted surgery, in the form of natural language expression of surgical scene understanding, can play a significant role in document entry ta…