7 papers · 1 filter
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
Alejandra Perez, Anita Rau, Lee White +4
Surgeons don't just see -- they interpret. When an expert observes a surgical scene, they understand not only what instrument is being used, but why it was chosen, what risk it pos…
SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
Alejandra Perez, Chinedu Nwoye, Ramtin Raji Kermani +2
Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without…
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
John J. Han, Adam Schmidt, Muhammad Abdullah Jamal +4
Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlo…
Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
Idris Hamoud, Vinkle Srivastav, Muhammad Abdullah Jamal +3
Understanding the workflow of surgical procedures in complex operating rooms requires a deep understanding of the interactions between clinicians and their environment. Surgical ac…
A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
Muhammad Abdullah Jamal, Omid Mohareri
In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Aut…
VidLPRO: A eo-anguage re-training Framework for botic and Laparoscopic Surgery
Mohammadmahdi Honarmand, Muhammad Abdullah Jamal, Omid Mohareri
We introduce VidLPRO, a novel video-language (VL) pre-training framework designed specifically for robotic and laparoscopic surgery. While existing surgical VL models primarily rel…