4 papers
DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models
John J. Han, Adam Schmidt, Muhammad Abdullah Jamal +2
Vision foundation models (VFMs) are valuable in data-scarce domains such as surgery, where a single pretrained backbone can provide rich representations for many downstream tasks.…
SCARED-C: Corrected Camera Poses for Endoscopic Depth Estimation
John J. Han, Adam Schmidt, Max Allan +2
The SCARED dataset is a widely used benchmark for endoscopic depth estimation, offering ground-truth 3D reconstructions captured with a structured light sensor. However, the depth…
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
John J. Han, Adam Schmidt, Muhammad Abdullah Jamal +4
Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlo…
Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
Danush Kumar Venkatesh, Adam Schmidt, Muhammad Abdullah Jamal +1
Surgical video datasets are essential for scene understanding, enabling procedural modeling and intra-operative support. However, these datasets are often heavily imbalanced, with…