papers

Publications (127)

cs.CV2026

Patient-Specific Articulated Digital Twins from a Single Full-Body CT Scan

Han Zhang, Boyang Zhao, Mathias Unberath

Patient-specific anatomical models provide individualized context for surgical planning, image-guided intervention, and algorithm development. However, most CT-derived models are s…

cs.CV2025

Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models

Yiqing Shen, Chenxiao Fan, Chenjia Li +1

The goal of text-to-video retrieval is to search large databases for relevant videos based on text queries. Existing methods have progressed to handling explicit queries where the…

cs.CV2018

Self-supervised Learning for Dense Depth Estimation in Monocular Endoscopy

Xingtong Liu, Ayushi Sinha, Mathias Unberath +4

We present a self-supervised approach to training convolutional neural networks for dense depth estimation from monocular endoscopy data without a priori modeling of anatomy or sha…

cs.CV2025

Privacy-Preserving Operating Room Workflow Analysis using Digital Twins

Alejandra Perez, Han Zhang, Yu-Chun Ku +7

The operating room (OR) is a complex environment where optimizing workflows is critical to reduce costs and improve patient outcomes. While computer vision approaches for automatic…

cs.HC2024

Human-AI collaboration is not very collaborative yet: A taxonomy of interaction patterns in AI-assisted decision making from a systematic review

Catalina Gomez, Sue Min Cho, Shichang Ke +2

Leveraging Artificial Intelligence (AI) in decision support systems has disproportionately focused on technological advancements, often overlooking the alignment between algorithmi…

cs.CV2018

Learning to See Forces: Surgical Force Prediction with RGB-Point Cloud Temporal Convolutional Networks

Cong Gao, Xingtong Liu, Michael Peven +2

Robotic surgery has been proven to offer clear advantages during surgical procedures, however, one of the major limitations is obtaining haptic feedback. Since it is often challeng…

cs.HC2020

Augment Yourself: Mixed Reality Self-Augmentation Using Optical See-through Head-mounted Displays and Physical Mirrors

Mathias Unberath, Kevin Yu, Roghayeh Barmaki +2

Optical see-though head-mounted displays (OST HMDs) are one of the key technologies for merging virtual objects and physical scenes to provide an immersive mixed reality (MR) envir…

cs.CV2025

Text-Driven Reasoning Video Editing via Reinforcement Learning on Digital Twin Representations

Yiqing Shen, Chenjia Li, Mathias Unberath

Text-driven video editing enables users to modify video content only using text queries. While existing methods can modify video content if explicit descriptions of editing targets…

eess.IV2021

An Interpretable Approach to Automated Severity Scoring in Pelvic Trauma

Anna Zapaishchykova, David Dreizin, Zhaoshuo Li +3

Pelvic ring disruptions result from blunt injury mechanisms and are often found in patients with multi-system trauma. To grade pelvic fracture severity in trauma victims based on w…

cs.CV2018

Augmented Reality-based Feedback for Technician-in-the-loop C-arm Repositioning

Mathias Unberath, Javad Fotouhi, Jonas Hajek +5

Interventional C-arm imaging is crucial to percutaneous orthopedic procedures as it enables the surgeon to monitor the progress of surgery on the anatomy level. Minimally invasive…

cs.CV2020

Reconstructing Sinus Anatomy from Endoscopic Video -- Towards a Radiation-free Approach for Quantitative Longitudinal Assessment

Xingtong Liu, Maia Stiber, Jindan Huang +4

Reconstructing accurate 3D surface models of sinus anatomy directly from an endoscopic video is a promising avenue for cross-sectional and longitudinal analysis to better understan…

cs.CV2018

On-the-fly Augmented Reality for Orthopaedic Surgery Using a Multi-Modal Fiducial

Sebastian Andress, Alex Johnson, Mathias Unberath +6

Fluoroscopic X-ray guidance is a cornerstone for percutaneous orthopaedic surgical procedures. However, two-dimensional observations of the three-dimensional anatomy suffer from th…

cs.CV2020

Multimodal and self-supervised representation learning for automatic gesture recognition in surgical robotics

Aniruddha Tamhane, Jie Ying Wu, Mathias Unberath

Self-supervised, multi-modal learning has been successful in holistic representation of complex scenarios. This can be useful to consolidate information from multiple modalities wh…

cs.CV2025

A Causal Framework for Aligning Image Quality Metrics and Deep Neural Network Robustness

Nathan Drenkow, Mathias Unberath

Image quality plays an important role in the performance of deep neural networks (DNNs) that have been widely shown to exhibit sensitivity to changes in imaging conditions. Convent…

cs.CV2024

An Endoscopic Chisel: Intraoperative Imaging Carves 3D Anatomical Models

Jan Emily Mangulabnan, Roger D. Soberanis-Mukul, Timo Teufel +7

Purpose: Preoperative imaging plays a pivotal role in sinus surgery where CTs offer patient-specific insights of complex anatomy, enabling real-time intraoperative navigation to co…

cs.CV2022

A Systematic Review of Robustness in Deep Learning for Computer Vision: Mind the gap?

Nathan Drenkow, Numair Sani, Ilya Shpitser +1

Deep neural networks for computer vision are deployed in increasingly safety-critical and socially-impactful applications, motivating the need to close the gap in model performance…

cs.CV2025

Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning

Yiqing Shen, Mathias Unberath

Visual reasoning may require models to interpret images and videos and respond to implicit text queries across diverse output formats, from pixel-level segmentation masks to natura…

cs.CV2023

RobustCLEVR: A Benchmark and Framework for Evaluating Robustness in Object-centric Learning

Nathan Drenkow, Mathias Unberath

Object-centric representation learning offers the potential to overcome limitations of image-level representations by explicitly parsing image scenes into their constituent compone…

cs.CV2026

AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction

Aiza Maksutova, Lalithkumar Seenivasan, Hao Ding +5

Surgical action automation has progressed rapidly toward achieving surgeon-like dexterous control, driven primarily by advances in learning from demonstration and vision-language-a…

physics.med-ph2018

Double Your Views - Exploiting Symmetry in Transmission Imaging

Alexander Preuhs, Andreas Maier, Michael Manhart +3

For a plane symmetric object we can find two views - mirrored at the plane of symmetry - that will yield the exact same image of that object. In consequence, having one image of a…

cs.HC2025

Human-AI Collaboration and Explainability for 2D/3D Registration Quality Assurance

Sue Min Cho, Alexander Do, Russell H. Taylor +1

Purpose: As surgery increasingly integrates advanced imaging, algorithms, and robotics to automate complex tasks, human judgment of system correctness remains a vital safeguard for…

cs.CV2017

UI-Net: Interactive Artificial Neural Networks for Iterative Image Segmentation Based on a User Model

Mario Amrehn, Sven Gaube, Mathias Unberath +6

For complex segmentation tasks, fully automatic systems are inherently limited in their achievable accuracy for extracting relevant objects. Especially in cases where only few data…

cs.RO2022

Rethinking Causality-driven Robot Tool Segmentation with Temporal Constraints

Hao Ding, Jie Ying Wu, Zhaoshuo Li +1

Purpose: Vision-based robot tool segmentation plays a fundamental role in surgical robots and downstream tasks. CaRTS, based on a complementary causal model, has shown promising pe…

cs.CV2022

Context-Enhanced Stereo Transformer

Weiyu Guo, Zhaoshuo Li, Yongkui Yang +5

Stereo depth estimation is of great interest for computer vision research. However, existing methods struggles to generalize and predict reliably in hazardous regions, such as larg…

cs.CV2025

Counterfactual World Models via Digital Twin-conditioned Video Diffusion

Yiqing Shen, Aiza Maksutova, Chenjia Li +1

World models learn to predict the temporal evolution of visual observations given a control signal, potentially enabling agents to reason about environments through forward simulat…

cs.CV2018

Plan in 2D, execute in 3D: An augmented reality solution for cup placement in total hip arthroplasty

Javad Fotouhi, Clayton P. Alexander, Mathias Unberath +9

Reproducibly achieving proper implant alignment is a critical step in total hip arthroplasty (THA) procedures that has been shown to substantially affect patient outcome. In curren…

cs.CV2025

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction

Yiqing Shen, Chenjia Li, Chenxiao Fan +1

Conventional approaches to video segmentation are confined to predefined object categories and cannot identify out-of-vocabulary objects, let alone objects that are not identified…

cs.RO2024

StraightTrack: Towards Mixed Reality Navigation System for Percutaneous K-wire Insertion

Han Zhang, Benjamin D. Killeen, Yu-Chun Ku +9

In percutaneous pelvic trauma surgery, accurate placement of Kirschner wires (K-wires) is crucial to ensure effective fracture fixation and avoid complications due to breaching the…

eess.IV2024

Promptable Counterfactual Diffusion Model for Unified Brain Tumor Segmentation and Generation with MRIs

Yiqing Shen, Guannan He, Mathias Unberath

Brain tumor analysis in Magnetic Resonance Imaging (MRI) is crucial for accurate diagnosis and treatment planning. However, the task remains challenging due to the complexity and v…

cs.CV2018

X-ray-transform Invariant Anatomical Landmark Detection for Pelvic Trauma Surgery

Bastian Bier, Mathias Unberath, Jan-Nico Zaech +5

X-ray image guidance enables percutaneous alternatives to complex procedures. Unfortunately, the indirect view onto the anatomy in addition to projective simplification substantial…

cs.RO2026

Humanoid Robots as First Assistants in Endoscopic Surgery

Sue Min Cho, Jan Emily Mangulabnan, Han Zhang +7

Humanoid robots have become a focal point of technological ambition, with claims of surgical capability within years in mainstream discourse. These projections are aspirational yet…

cs.HC2022

Explainable Medical Imaging AI Needs Human-Centered Design: Guidelines and Evidence from a Systematic Review

Haomin Chen, Catalina Gomez, Chien-Ming Huang +1

Transparency in Machine Learning (ML), attempts to reveal the working mechanisms of complex models. Transparent ML promises to advance human factors engineering goals of human-cent…

cs.CV2025

Towards Robust Algorithms for Surgical Phase Recognition via Digital Twin Representation

Hao Ding, Yuqian Zhang, Wenzheng Cheng +7

Surgical phase recognition (SPR) is an integral component of surgical data science, enabling high-level surgical analysis. End-to-end trained neural networks that predict surgical…

physics.med-ph2018

DeepDRR -- A Catalyst for Machine Learning in Fluoroscopy-guided Procedures

Mathias Unberath, Jan-Nico Zaech, Sing Chun Lee +4

Machine learning-based approaches outperform competing methods in most disciplines relevant to diagnostic radiology. Interventional radiology, however, has not yet benefited substa…

cs.CV2020

Generalizing Spatial Transformers to Projective Geometry with Applications to 2D/3D Registration

Cong Gao, Xingtong Liu, Wenhao Gu +4

Differentiable rendering is a technique to connect 3D scenes with corresponding 2D images. Since it is differentiable, processes during image formation can be learned. Previous app…

cs.HC2024

Explainable AI Enhances Glaucoma Referrals, Yet the Human-AI Team Still Falls Short of the AI Alone

Catalina Gomez, Ruolin Wang, Katharina Breininger +6

Primary care providers are vital for initial triage and referrals to specialty care. In glaucoma, asymptomatic and fast progression can lead to vision loss, necessitating timely re…

cs.CV2019

Self-supervised Dense 3D Reconstruction from Monocular Endoscopic Video

Xingtong Liu, Ayushi Sinha, Masaru Ishii +3

We present a self-supervised learning-based pipeline for dense 3D reconstruction from full-length monocular endoscopic videos without a priori modeling of anatomy or shading. Our m…

cs.CV2026

TwinOR: Photorealistic Digital Twins of Dynamic Operating Rooms for Embodied AI Research

Han Zhang, Yiqing Shen, Roger D. Soberanis-Mukul +11

Developing embodied AI for intelligent surgical systems requires safe, controllable environments for continual learning and evaluation. However, safety regulations and operational…

cs.OH2017

Technical Note: Towards Virtual Monitors for Image Guided Interventions - Real-time Streaming to Optical See-Through Head-Mounted Displays

Long Qian, Mathias Unberath, Kevin Yu +4

Purpose: Image guidance is crucial for the success of many interventions. Images are displayed on designated monitors that cannot be positioned optimally due to sterility and spati…

cs.LG2023

Pelphix: Surgical Phase Recognition from X-ray Images in Percutaneous Pelvic Fixation

Benjamin D. Killeen, Han Zhang, Jan Mangulabnan +4

Surgical phase recognition (SPR) is a crucial element in the digital transformation of the modern operating theater. While SPR based on video sources is well-established, incorpora…

cs.RO2025

Constrained Natural Language Action Planning for Resilient Embodied Systems

Grayson Byrd, Corban Rivera, Bethany Kemp +5

Replicating human-level intelligence in the execution of embodied tasks remains challenging due to the unconstrained nature of real-world environments. Novel use of large language…

eess.IV2021

Pose-dependent weights and Domain Randomization for fully automatic X-ray to CT Registration

Matthias Grimm, Javier Esteban, Mathias Unberath +1

Fully automatic X-ray to CT registration requires a solid initialization to provide an initial alignment within the capture range of existing intensity-based registrations. This wo…

eess.IV2024

FastSAM-3DSlicer: A 3D-Slicer Extension for 3D Volumetric Segment Anything Model with Uncertainty Quantification

Yiqing Shen, Xinyuan Shao, Blanca Inigo Romillo +2

Accurate segmentation of anatomical structures and pathological regions in medical images is crucial for diagnosis, treatment planning, and disease monitoring. While the Segment An…

cs.CV2026

SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation

Sampath Rapuri, Lalithkumar Seenivasan, Dominik Schneider +9

A surgical world model capable of generating realistic surgical action videos with precise control over tool-tissue interactions can address fundamental challenges in surgical AI a…

cs.CV2022

SAGE: SLAM with Appearance and Geometry Prior for Endoscopy

Xingtong Liu, Zhaoshuo Li, Masaru Ishii +3

In endoscopy, many applications (e.g., surgical navigation) would benefit from a real-time method that can simultaneously track the endoscope and reconstruct the dense 3D geometry…

eess.IV2025

Benchmark of Segmentation Techniques for Pelvic Fracture in CT and X-ray: Summary of the PENGWIN 2024 Challenge

Yudi Sang, Yanzhen Liu, Sutuke Yibulayimu +33

The segmentation of pelvic fracture fragments in CT and X-ray images is crucial for trauma diagnosis, surgical planning, and intraoperative guidance. However, accurately and effici…

cs.CV2025

Causality-Driven Audits of Model Robustness

Nathan Drenkow, William Paul, Chris Ribaudo +1

Robustness audits of deep neural networks (DNN) provide a means to uncover model sensitivities to the challenging real-world imaging conditions that significantly degrade DNN perfo…

cs.CV2025

Fast Reasoning Segmentation for Images and Videos

Yiqing Shen, Mathias Unberath

Reasoning segmentation enables open-set object segmentation via implicit text queries, therefore serving as a foundation for embodied agents that should operate autonomously in rea…

cs.CV2018

Closing the Calibration Loop: An Inside-out-tracking Paradigm for Augmented Reality in Orthopedic Surgery

Jonas Hajek, Mathias Unberath, Javad Fotouhi +6

In percutaneous orthopedic interventions the surgeon attempts to reduce and fixate fractures in bony structures. The complexity of these interventions arises when the surgeon perfo…

eess.IV2025

Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins

Yiqing Shen, Chenjia Li, Bohan Liu +3

Analyzing operating room (OR) workflows to derive quantitative insights into OR efficiency is important for hospitals to maximize patient care and financial sustainability. Prior w…

cs.CV2021

On the Sins of Image Synthesis Loss for Self-supervised Depth Estimation

Zhaoshuo Li, Nathan Drenkow, Hao Ding +5

Scene depth estimation from stereo and monocular imagery is critical for extracting 3D information for downstream tasks such as scene understanding. Recently, learning-based method…

cs.RO2020

Leveraging Vision and Kinematics Data to Improve Realism of Biomechanic Soft-tissue Simulation for Robotic Surgery

Jie Ying Wu, Peter Kazanzides, Mathias Unberath

Purpose Surgical simulations play an increasingly important role in surgeon education and developing algorithms that enable robots to perform surgical subtasks. To model anatomy, F…

cs.CV2021

E-DSSR: Efficient Dynamic Surgical Scene Reconstruction with Transformer-based Stereoscopic Depth Perception

Yonghao Long, Zhaoshuo Li, Chi Hang Yee +4

Reconstructing the scene of robotic surgery from the stereo endoscopic video is an important and promising topic in surgical data science, which potentially supports many applicati…

cs.CV2021

The Impact of Machine Learning on 2D/3D Registration for Image-guided Interventions: A Systematic Review and Perspective

Mathias Unberath, Cong Gao, Yicheng Hu +4

Image-based navigation is widely considered the next frontier of minimally invasive surgery. It is believed that image-based navigation will increase the access to reproducible, sa…

cs.CV2023

A Quantitative Evaluation of Dense 3D Reconstruction of Sinus Anatomy from Monocular Endoscopic Video

Jan Emily Mangulabnan, Roger D. Soberanis-Mukul, Timo Teufel +9

Generating accurate 3D reconstructions from endoscopic video is a promising avenue for longitudinal radiation-free analysis of sinus anatomy and surgical outcomes. Several methods…

cs.HC2017

Seminar Innovation Management - Winter Term 2017

Gerd Häusler, Aleksandra Milczarek, Markus Schreiter +14

This document contains the results obtained by the Innovation Management Seminar in winter term 2017. In total 11 ideas have been developed by the team. In the document all 11 idea…

cs.RO2024

Intelligent Control of Robotic X-ray Devices using a Language-promptable Digital Twin

Benjamin D. Killeen, Anushri Suresh, Catalina Gomez +3

Natural language offers a convenient, flexible interface for controlling robotic C-arm X-ray systems, making advanced functionality and controls accessible. However, enabling langu…

cs.CV2023

Neuralangelo: High-Fidelity Neural Surface Reconstruction

Zhaoshuo Li, Thomas Müller, Alex Evans +4

Neural surface reconstruction has been shown to be powerful for recovering dense 3D surfaces via image-based neural rendering. However, current methods struggle to recover detailed…

cs.CV2021

Pediatric Otoscopy Video Screening with Shift Contrastive Anomaly Detection

Weiyao Wang, Aniruddha Tamhane, Christine Santos +4

Ear related concerns and symptoms represents the leading indication for seeking pediatric healthcare attention. Despite the high incidence of such encounters, the diagnostic proces…

cs.LG2024

An Intrinsically Explainable Approach to Detecting Vertebral Compression Fractures in CT Scans via Neurosymbolic Modeling

Blanca Inigo, Yiqing Shen, Benjamin D. Killeen +4

Vertebral compression fractures (VCFs) are a common and potentially serious consequence of osteoporosis. Yet, they often remain undiagnosed. Opportunistic screening, which involves…

eess.IV2025

Neural Finite-State Machines for Surgical Phase Recognition

Hao Ding, Zhongpai Gao, Benjamin Planche +7

Surgical phase recognition (SPR) is crucial for applications in workflow optimization, performance evaluation, and real-time intervention guidance. However, current deep learning m…

cs.CV2020

Fast and Automatic Periacetabular Osteotomy Fragment Pose Estimation Using Intraoperatively Implanted Fiducials and Single-View Fluoroscopy

Robert Grupp, Ryan Murphy, Rachel Hegeman +6

Accurate and consistent mental interpretation of fluoroscopy to determine the position and orientation of acetabular bone fragments in 3D space is difficult. We propose a computer…

cs.CV2023

Task-based Generation of Optimized Projection Sets using Differentiable Ranking

Linda-Sophie Schneider, Mareike Thies, Christopher Syben +3

We present a method for selecting valuable projections in computed tomography (CT) scans to enhance image reconstruction and diagnosis. The approach integrates two important factor…

cs.CV2026

Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA

Yiqing Shen, Han Zhang, Mathias Unberath

Surgical video question answering requires multi-step reasoning across semantic, spatial, and temporal dimensions. Existing methods architecturally compress videos into discrete to…

cs.CV2020

Exploring Partial Intrinsic and Extrinsic Symmetry in 3D Medical Imaging

Javad Fotouhi, Giacomo Taylor, Mathias Unberath +5

We present a novel methodology to detect imperfect bilateral symmetry in CT of human anatomy. In this paper, the structurally symmetric nature of the pelvic bone is explored and is…

q-bio.TO2019

A Biomechanical Study on the Use of Curved Drilling Technique for Treatment of Osteonecrosis of Femoral Head

Mahsan Bakhtiarinejad, Farshid Alambeigi, Alireza Chamani +3

Osteonecrosis occurs due to the loss of blood supply to the bone, leading to spontaneous death of the trabecular bone. Delayed treatment of the involved patients results in collaps…

cs.CV2022

Temporally Consistent Online Depth Estimation in Dynamic Scenes

Zhaoshuo Li, Wei Ye, Dilin Wang +4

Temporally consistent depth estimation is crucial for online applications such as augmented reality. While stereo depth estimation has received substantial attention as a promising…

cs.CV2025

RVTBench: A Benchmark for Visual Reasoning Tasks

Yiqing Shen, Chenjia Li, Chenxiao Fan +1

Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the…

eess.IV2026

2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Models

Sampath Rapuri, Jeremy Ko, Benjamin D. Killeen +2

The ability to synthesize realistic X-ray images has catalyzed the development of AI models for X-ray image-guided procedures, which otherwise suffer from a lack of available annot…

cs.CY2020

A County-level Dataset for Informing the United States' Response to COVID-19

Benjamin D. Killeen, Jie Ying Wu, Kinjal Shah +8

As the coronavirus disease 2019 (COVID-19) continues to be a global pandemic, policy makers have enacted and reversed non-pharmaceutical interventions with various levels of restri…

cs.CV2023

TAToo: Vision-based Joint Tracking of Anatomy and Tool for Skull-base Surgery

Zhaoshuo Li, Hongchao Shu, Ruixing Liang +5

Purpose: Tracking the 3D motion of the surgical tool and the patient anatomy is a fundamental requirement for computer-assisted skull-base surgery. The estimated motion can be used…

cs.RO2021

Virtual Reality for Synergistic Surgical Training and Data Generation

Adnan Munawar, Zhaoshuo Li, Punit Kunjam +7

Surgical simulators not only allow planning and training of complex procedures, but also offer the ability to generate structured data for algorithm development, which may be appli…

eess.IV2019

CAI4CAI: The Rise of Contextual Artificial Intelligence in Computer Assisted Interventions

Tom Vercauteren, Mathias Unberath, Nicolas Padoy +1

Data-driven computational approaches have evolved to enable extraction of information from medical images with a reliability, accuracy and speed which is already transforming their…

eess.IV2019

Learning to Avoid Poor Images: Towards Task-aware C-arm Cone-beam CT Trajectories

Jan-Nico Zaech, Cong Gao, Bastian Bier +4

Metal artifacts in computed tomography (CT) arise from a mismatch between physics of image formation and idealized assumptions during tomographic reconstruction. These artifacts ar…

cs.LG2023

Data AUDIT: Identifying Attribute Utility- and Detectability-Induced Bias in Task Models

Mitchell Pavlak, Nathan Drenkow, Nicholas Petrick +2

To safely deploy deep learning-based computer vision models for computer-aided detection and diagnosis, we must ensure that they are robust and reliable. Towards that goal, algorit…

cs.RO2022

CaRTS: Causality-driven Robot Tool Segmentation from Vision and Kinematics Data

Hao Ding, Jintan Zhang, Peter Kazanzides +2

Vision-based segmentation of the robotic tool during robot-assisted surgery enables downstream applications, such as augmented reality feedback, while allowing for inaccuracies in…

cs.CV2020

Automatic Annotation of Hip Anatomy in Fluoroscopy for Robust and Efficient 2D/3D Registration

Robert Grupp, Mathias Unberath, Cong Gao +7

Fluoroscopy is the standard imaging modality used to guide hip surgery and is therefore a natural sensor for computer-assisted navigation. In order to efficiently solve the complex…

cs.CV2020

A Learning-based Method for Online Adjustment of C-arm Cone-Beam CT Source Trajectories for Artifact Avoidance

Mareike Thies, Jan-Nico Zäch, Cong Gao +4

During spinal fusion surgery, screws are placed close to critical nerves suggesting the need for highly accurate screw placement. Verifying screw placement on high-quality tomograp…

cs.CV2026

Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins

Yiqing Shen, Hao Ding, Mathias Unberath

Text-to-video retrieval in operating rooms (OR) is an enabling technology for OR safety, as it allows stakeholders to retrieve and inspect recordings of specific events. However, b…

eess.IV2020

From Perspective X-ray Imaging to Parallax-Robust Orthographic Stitching

Javad Fotouhi, Xingtong Liu, Mehran Armand +2

Stitching images acquired under perspective projective geometry is a relevant topic in computer vision with multiple applications ranging from smartphone panoramas to the construct…

cs.CV2025

Did you just see that? Arbitrary view synthesis for egocentric replay of operating room workflows from ambient sensors

Han Zhang, Lalithkumar Seenivasan, Jose L. Porras +9

Observing surgical practice has historically relied on fixed vantage points or recollections, leaving the egocentric visual perspectives that guide clinical decisions undocumented.…

cs.CV2021

An Interpretable Algorithm for Uveal Melanoma Subtyping from Whole Slide Cytology Images

Haomin Chen, T. Y. Alvin Liu, Catalina Gomez +2

Algorithmic decision support is rapidly becoming a staple of personalized medicine, especially for high-stakes recommendations in which access to certain information can drasticall…

cs.CV2023

Automated Artifact Detection in Ultra-widefield Fundus Photography of Patients with Sickle Cell Disease

Anqi Feng, Dimitri Johnson, Grace R. Reilly +5

Importance: Ultra-widefield fundus photography (UWF-FP) has shown utility in sickle cell retinopathy screening; however, image artifact may diminish quality and gradeability of ima…

cs.CV2025

Reasoning Segmentation for Images and Videos: A Survey

Yiqing Shen, Chenjia Li, Fei Xiong +4

Reasoning Segmentation (RS) aims to delineate objects based on implicit text queries, the interpretation of which requires reasoning and knowledge integration. Unlike the tradition…

cs.CV2026

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

Chenyan Jing, Hao Ding, Lalithkumar Seenivasan +2

Vision foundation models such as SAM 3 can provide transferable object-level structure across diverse surgical video conditions, but segmentation outputs do not explicitly encode t…

cs.CV2026

Towards Controllable Video Synthesis of Routine and Rare OR Events

Dominik Schneider, Lalithkumar Seenivasan, Sampath Rapuri +8

Purpose: Curating large-scale datasets of operating room (OR) workflow, encompassing rare, safety-critical, or atypical events, remains operationally and ethically challenging. Thi…

cs.CV2025

FluoroSAM: A Language-promptable Foundation Model for Flexible X-ray Image Segmentation

Benjamin D. Killeen, Liam J. Wang, Blanca Inigo +5

Language promptable X-ray image segmentation would enable greater flexibility for human-in-the-loop workflows in diagnostic and interventional precision medicine. Prior efforts hav…

cs.CV2024

From Generalization to Precision: Exploring SAM for Tool Segmentation in Surgical Environments

Kanyifeechukwu J. Oguine, Roger D. Soberanis-Mukul, Nathan Drenkow +1

Purpose: Accurate tool segmentation is essential in computer-aided procedures. However, this task conveys challenges due to artifacts' presence and the limited training data in med…

cs.HC2025

Explainable AI for Automated User-specific Feedback in Surgical Skill Acquisition

Catalina Gomez, Lalithkumar Seenivasan, Xinrui Zou +9

Traditional surgical skill acquisition relies heavily on expert feedback, yet direct access is limited by faculty availability and variability in subjective assessments. While trai…

cs.AI2024

ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and Execution

Corban Rivera, Grayson Byrd, William Paul +12

Robotic planning and execution in open-world environments is a complex problem due to the vast state spaces and high variability of task embodiment. Recent advances in perception a…

eess.IV2024

Performance and Non-adversarial Robustness of the Segment Anything Model 2 in Surgical Video Segmentation

Yiqing Shen, Hao Ding, Xinyuan Shao +1

Fully supervised deep learning (DL) models for surgical video segmentation have been shown to struggle with non-adversarial, real-world corruptions of image quality including smoke…

cs.CV2026

Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation

Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding +5

Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation acr…

cs.CV2021

Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with Transformers

Zhaoshuo Li, Xingtong Liu, Nathan Drenkow +4

Stereo depth estimation relies on optimal correspondence matching between pixels on epipolar lines in the left and right images to infer depth. In this work, we revisit the problem…

cs.CV2018

Exploiting Partial Structural Symmetry For Patient-Specific Image Augmentation in Trauma Interventions

Javad Fotouhi, Mathias Unberath, Giacomo Taylor +7

In unilateral pelvic fracture reductions, surgeons attempt to reconstruct the bone fragments such that bilateral symmetry in the bony anatomy is restored. We propose to exploit thi…

cs.CV2025

BronchOpt : Vision-Based Pose Optimization with Fine-Tuned Foundation Models for Accurate Bronchoscopy Navigation

Hongchao Shu, Roger D. Soberanis-Mukul, Jiru Xu +6

Accurate intra-operative localization of the bronchoscope tip relative to patient anatomy remains challenging due to respiratory motion, anatomical variability, and CT-to-body dive…

cs.RO2025

Beyond Rigid AI: Towards Natural Human-Machine Symbiosis for Interoperative Surgical Assistance

Lalithkumar Seenivasan, Jiru Xu, Roger D. Soberanis Mukul +6

Emerging surgical data science and robotics solutions, especially those designed to provide assistance in situ, require natural human-machine interfaces to fully unlock their poten…

cs.RO2020

Reflective-AR Display: An Interaction Methodology for Virtual-Real Alignment in Medical Robotics

Javad Fotouhi, Tianyu Song, Arian Mehrfard +8

Robot-assisted minimally invasive surgery has shown to improve patient outcomes, as well as reduce complications and recovery time for several clinical applications. While increasi…

eess.IV2023

TransNuSeg: A Lightweight Multi-Task Transformer for Nuclei Segmentation

Zhenqi He, Mathias Unberath, Jing Ke +1

Nuclei appear small in size, yet, in real clinical practice, the global spatial information and correlation of the color or brightness contrast between nuclei and background, have…

eess.IV2022

SyntheX: Scaling Up Learning-based X-ray Image Analysis Through In Silico Experiments

Cong Gao, Benjamin D. Killeen, Yicheng Hu +4

Artificial intelligence (AI) now enables automated interpretation of medical images for clinical use. However, AI's potential use for interventional images (versus those involved i…

cs.RO2026

Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models

Hao Ding, Lalithkumar Seenivasan, Hongchao Shu +7

Large language model-based (LLM) agents are emerging as a powerful enabler of robust embodied intelligence due to their capability of planning complex action sequences. Sound plann…