18 papers
2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Models
Sampath Rapuri, Jeremy Ko, Benjamin D. Killeen +2
The ability to synthesize realistic X-ray images has catalyzed the development of AI models for X-ray image-guided procedures, which otherwise suffer from a lack of available annot…
Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins
Yiqing Shen, Hao Ding, Mathias Unberath
Text-to-video retrieval in operating rooms (OR) is an enabling technology for OR safety, as it allows stakeholders to retrieve and inspect recordings of specific events. However, b…
Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA
Yiqing Shen, Han Zhang, Mathias Unberath
Surgical video question answering requires multi-step reasoning across semantic, spatial, and temporal dimensions. Existing methods architecturally compress videos into discrete to…
TwinOR: Photorealistic Digital Twins of Dynamic Operating Rooms for Embodied AI Research
Han Zhang, Yiqing Shen, Roger D. Soberanis-Mukul +11
Developing embodied AI for intelligent surgical systems requires safe, controllable environments for continual learning and evaluation. However, safety regulations and operational…
Towards Controllable Video Synthesis of Routine and Rare OR Events
Dominik Schneider, Lalithkumar Seenivasan, Sampath Rapuri +8
Purpose: Curating large-scale datasets of operating room (OR) workflow, encompassing rare, safety-critical, or atypical events, remains operationally and ethically challenging. Thi…
Counterfactual World Models via Digital Twin-conditioned Video Diffusion
Yiqing Shen, Aiza Maksutova, Chenjia Li +1
World models learn to predict the temporal evolution of visual observations given a control signal, potentially enabling agents to reason about environments through forward simulat…