20 papers
SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding
Yuqing Feng, Jiawei Ma, Kevin Qinghong Lin +6
Surgical procedures unfold as structured and recurring clinical events, whose real-time understanding via intraoperative surgical videos is critical for intraoperative decision-mak…
HyperVLP: Enhancing Hierarchical Surgical Video-Language Pre-training in Hyperbolic Space
Yaojun Hu, Kun Yuan, Nassir Navab +3
Surgical vision-language foundation models typically adopt educational materials, such as surgical lecture videos, to transfer surgical knowledge encoded in language into visual re…
Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025
Aneeq Zia, Max Berniker, Rogerio Garcia Nespolo +153
Robotic assisted (RA) surgery promises to transform surgical intervention. Intuitive Surgical is committed to fostering these changes and the machine learning models and algorithms…
CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
Florian Stilz, Vinkle Srivastav, Nassir Navab +1
Video-language foundation models have proven to be highly effective in zero-shot applications across a wide range of tasks. A particularly challenging area is the intraoperative su…
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
Kun Yuan, Min Woo Sun, Zhen Chen +7
There is a growing interest in developing strong biomedical vision-language models. A popular approach to achieve robust representations is to use web-scale scientific data. Howeve…
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
Xiangteng He, Shunsuke Sakai, Shivam Chandhok +5
Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for learning transferable features. In p…