10 papers
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture
Felix Tristram, Stefano Gasperini, Benjamin Killeen +4
The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assistance systems for complex, mu…
OR-Action: Multi-Role Video Understanding with Fine-Grained Actions
Felix Tristram, Ege Ãzsoy, Christian Benz +3
Fine-grained understanding of operating room (OR) activity could enable workflow-aware assistance, yet remains difficult due to clutter, occlusions, and limited sensing. The prevai…
SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation
Ssharvien Kumar Sivakumar, Akwele Johnson, Anirudh Dhingra +3
Realistic surgical simulation plays a crucial role in training novice surgeons and in the development of autonomous agents. World models can scale such simulation environments to r…
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
Nikolo Rohrmoser, Ghazal Ghazaei, Michael Sommersperger +1
Purpose: The integration of multimodal imaging into operating rooms paves the way for comprehensive surgical scene understanding. In ophthalmic surgery, by now, two complementary i…
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
Felix Holm, Ghazal Ghazaei, Nassir Navab
Purpose: Detailed surgical recognition is critical for advancing AI-assisted surgery, yet progress is hampered by high annotation costs, data scarcity, and a lack of interpretable…
CAT-SG: A Large Dynamic Scene Graph Dataset for Fine-Grained Understanding of Cataract Surgery
Felix Holm, Gözde Ãnver, Ghazal Ghazaei +1
Understanding the intricate workflows of cataract surgery requires modeling complex interactions between surgical tools, anatomical structures, and procedural techniques. Existing…