collaborators

5 papers

cs.CV2026

SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning

Alejandra Perez, Anita Rau, Lee White +4

Surgeons don't just see -- they interpret. When an expert observes a surgical scene, they understand not only what instrument is being used, but why it was chosen, what risk it pos…

cs.CV2026

SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning

Alejandra Perez, Chinedu Nwoye, Ramtin Raji Kermani +2

Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without…

cs.CV2026

On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training

John J. Han, Adam Schmidt, Muhammad Abdullah Jamal +4

Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlo…

q-bio.QM2025

Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model

Danush Kumar Venkatesh, Adam Schmidt, Muhammad Abdullah Jamal +1

Surgical video datasets are essential for scene understanding, enabling procedural modeling and intra-operative support. However, these datasets are often heavily imbalanced, with…

cs.CV2025

Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition

Idris Hamoud, Vinkle Srivastav, Muhammad Abdullah Jamal +3

Understanding the workflow of surgical procedures in complex operating rooms requires a deep understanding of the interactions between clinicians and their environment. Surgical ac…