5 papers
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
Alejandra Perez, Anita Rau, Lee White +4
Surgeons don't just see -- they interpret. When an expert observes a surgical scene, they understand not only what instrument is being used, but why it was chosen, what risk it pos…
SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
Alejandra Perez, Chinedu Nwoye, Ramtin Raji Kermani +2
Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without…
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
John J. Han, Adam Schmidt, Muhammad Abdullah Jamal +4
Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlo…
Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
Danush Kumar Venkatesh, Adam Schmidt, Muhammad Abdullah Jamal +1
Surgical video datasets are essential for scene understanding, enabling procedural modeling and intra-operative support. However, these datasets are often heavily imbalanced, with…
Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
Idris Hamoud, Vinkle Srivastav, Muhammad Abdullah Jamal +3
Understanding the workflow of surgical procedures in complex operating rooms requires a deep understanding of the interactions between clinicians and their environment. Surgical ac…