Publications (127)
Patient-Specific Articulated Digital Twins from a Single Full-Body CT Scan
Han Zhang, Boyang Zhao, Mathias Unberath
Patient-specific anatomical models provide individualized context for surgical planning, image-guided intervention, and algorithm development. However, most CT-derived models are s…
Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models
Yiqing Shen, Chenxiao Fan, Chenjia Li +1
The goal of text-to-video retrieval is to search large databases for relevant videos based on text queries. Existing methods have progressed to handling explicit queries where the…
Self-supervised Learning for Dense Depth Estimation in Monocular Endoscopy
Xingtong Liu, Ayushi Sinha, Mathias Unberath +4
We present a self-supervised approach to training convolutional neural networks for dense depth estimation from monocular endoscopy data without a priori modeling of anatomy or sha…
Privacy-Preserving Operating Room Workflow Analysis using Digital Twins
Alejandra Perez, Han Zhang, Yu-Chun Ku +7
The operating room (OR) is a complex environment where optimizing workflows is critical to reduce costs and improve patient outcomes. While computer vision approaches for automatic…
Human-AI collaboration is not very collaborative yet: A taxonomy of interaction patterns in AI-assisted decision making from a systematic review
Catalina Gomez, Sue Min Cho, Shichang Ke +2
Leveraging Artificial Intelligence (AI) in decision support systems has disproportionately focused on technological advancements, often overlooking the alignment between algorithmi…
Learning to See Forces: Surgical Force Prediction with RGB-Point Cloud Temporal Convolutional Networks
Cong Gao, Xingtong Liu, Michael Peven +2
Robotic surgery has been proven to offer clear advantages during surgical procedures, however, one of the major limitations is obtaining haptic feedback. Since it is often challeng…
Augment Yourself: Mixed Reality Self-Augmentation Using Optical See-through Head-mounted Displays and Physical Mirrors
Mathias Unberath, Kevin Yu, Roghayeh Barmaki +2
Optical see-though head-mounted displays (OST HMDs) are one of the key technologies for merging virtual objects and physical scenes to provide an immersive mixed reality (MR) envir…
Text-Driven Reasoning Video Editing via Reinforcement Learning on Digital Twin Representations
Yiqing Shen, Chenjia Li, Mathias Unberath
Text-driven video editing enables users to modify video content only using text queries. While existing methods can modify video content if explicit descriptions of editing targets…
An Interpretable Approach to Automated Severity Scoring in Pelvic Trauma
Anna Zapaishchykova, David Dreizin, Zhaoshuo Li +3
Pelvic ring disruptions result from blunt injury mechanisms and are often found in patients with multi-system trauma. To grade pelvic fracture severity in trauma victims based on w…
Augmented Reality-based Feedback for Technician-in-the-loop C-arm Repositioning
Mathias Unberath, Javad Fotouhi, Jonas Hajek +5
Interventional C-arm imaging is crucial to percutaneous orthopedic procedures as it enables the surgeon to monitor the progress of surgery on the anatomy level. Minimally invasive…
Reconstructing Sinus Anatomy from Endoscopic Video -- Towards a Radiation-free Approach for Quantitative Longitudinal Assessment
Xingtong Liu, Maia Stiber, Jindan Huang +4
Reconstructing accurate 3D surface models of sinus anatomy directly from an endoscopic video is a promising avenue for cross-sectional and longitudinal analysis to better understan…
On-the-fly Augmented Reality for Orthopaedic Surgery Using a Multi-Modal Fiducial
Sebastian Andress, Alex Johnson, Mathias Unberath +6
Fluoroscopic X-ray guidance is a cornerstone for percutaneous orthopaedic surgical procedures. However, two-dimensional observations of the three-dimensional anatomy suffer from th…
Multimodal and self-supervised representation learning for automatic gesture recognition in surgical robotics
Aniruddha Tamhane, Jie Ying Wu, Mathias Unberath
Self-supervised, multi-modal learning has been successful in holistic representation of complex scenarios. This can be useful to consolidate information from multiple modalities wh…
A Causal Framework for Aligning Image Quality Metrics and Deep Neural Network Robustness
Nathan Drenkow, Mathias Unberath
Image quality plays an important role in the performance of deep neural networks (DNNs) that have been widely shown to exhibit sensitivity to changes in imaging conditions. Convent…
An Endoscopic Chisel: Intraoperative Imaging Carves 3D Anatomical Models
Jan Emily Mangulabnan, Roger D. Soberanis-Mukul, Timo Teufel +7
Purpose: Preoperative imaging plays a pivotal role in sinus surgery where CTs offer patient-specific insights of complex anatomy, enabling real-time intraoperative navigation to co…
A Systematic Review of Robustness in Deep Learning for Computer Vision: Mind the gap?
Nathan Drenkow, Numair Sani, Ilya Shpitser +1
Deep neural networks for computer vision are deployed in increasingly safety-critical and socially-impactful applications, motivating the need to close the gap in model performance…
Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
Yiqing Shen, Mathias Unberath
Visual reasoning may require models to interpret images and videos and respond to implicit text queries across diverse output formats, from pixel-level segmentation masks to natura…
RobustCLEVR: A Benchmark and Framework for Evaluating Robustness in Object-centric Learning
Nathan Drenkow, Mathias Unberath
Object-centric representation learning offers the potential to overcome limitations of image-level representations by explicitly parsing image scenes into their constituent compone…
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
Aiza Maksutova, Lalithkumar Seenivasan, Hao Ding +5
Surgical action automation has progressed rapidly toward achieving surgeon-like dexterous control, driven primarily by advances in learning from demonstration and vision-language-a…
Double Your Views - Exploiting Symmetry in Transmission Imaging
Alexander Preuhs, Andreas Maier, Michael Manhart +3
For a plane symmetric object we can find two views - mirrored at the plane of symmetry - that will yield the exact same image of that object. In consequence, having one image of a…
Human-AI Collaboration and Explainability for 2D/3D Registration Quality Assurance
Sue Min Cho, Alexander Do, Russell H. Taylor +1
Purpose: As surgery increasingly integrates advanced imaging, algorithms, and robotics to automate complex tasks, human judgment of system correctness remains a vital safeguard for…
UI-Net: Interactive Artificial Neural Networks for Iterative Image Segmentation Based on a User Model
Mario Amrehn, Sven Gaube, Mathias Unberath +6
For complex segmentation tasks, fully automatic systems are inherently limited in their achievable accuracy for extracting relevant objects. Especially in cases where only few data…
Rethinking Causality-driven Robot Tool Segmentation with Temporal Constraints
Hao Ding, Jie Ying Wu, Zhaoshuo Li +1
Purpose: Vision-based robot tool segmentation plays a fundamental role in surgical robots and downstream tasks. CaRTS, based on a complementary causal model, has shown promising pe…
Context-Enhanced Stereo Transformer
Weiyu Guo, Zhaoshuo Li, Yongkui Yang +5
Stereo depth estimation is of great interest for computer vision research. However, existing methods struggles to generalize and predict reliably in hazardous regions, such as larg…
Counterfactual World Models via Digital Twin-conditioned Video Diffusion
Yiqing Shen, Aiza Maksutova, Chenjia Li +1
World models learn to predict the temporal evolution of visual observations given a control signal, potentially enabling agents to reason about environments through forward simulat…
Plan in 2D, execute in 3D: An augmented reality solution for cup placement in total hip arthroplasty
Javad Fotouhi, Clayton P. Alexander, Mathias Unberath +9
Reproducibly achieving proper implant alignment is a critical step in total hip arthroplasty (THA) procedures that has been shown to substantially affect patient outcome. In curren…
Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction
Yiqing Shen, Chenjia Li, Chenxiao Fan +1
Conventional approaches to video segmentation are confined to predefined object categories and cannot identify out-of-vocabulary objects, let alone objects that are not identified…
StraightTrack: Towards Mixed Reality Navigation System for Percutaneous K-wire Insertion
Han Zhang, Benjamin D. Killeen, Yu-Chun Ku +9
In percutaneous pelvic trauma surgery, accurate placement of Kirschner wires (K-wires) is crucial to ensure effective fracture fixation and avoid complications due to breaching the…
Promptable Counterfactual Diffusion Model for Unified Brain Tumor Segmentation and Generation with MRIs
Yiqing Shen, Guannan He, Mathias Unberath
Brain tumor analysis in Magnetic Resonance Imaging (MRI) is crucial for accurate diagnosis and treatment planning. However, the task remains challenging due to the complexity and v…
X-ray-transform Invariant Anatomical Landmark Detection for Pelvic Trauma Surgery
Bastian Bier, Mathias Unberath, Jan-Nico Zaech +5
X-ray image guidance enables percutaneous alternatives to complex procedures. Unfortunately, the indirect view onto the anatomy in addition to projective simplification substantial…
Humanoid Robots as First Assistants in Endoscopic Surgery
Sue Min Cho, Jan Emily Mangulabnan, Han Zhang +7
Humanoid robots have become a focal point of technological ambition, with claims of surgical capability within years in mainstream discourse. These projections are aspirational yet…
Explainable Medical Imaging AI Needs Human-Centered Design: Guidelines and Evidence from a Systematic Review
Haomin Chen, Catalina Gomez, Chien-Ming Huang +1
Transparency in Machine Learning (ML), attempts to reveal the working mechanisms of complex models. Transparent ML promises to advance human factors engineering goals of human-cent…
Towards Robust Algorithms for Surgical Phase Recognition via Digital Twin Representation
Hao Ding, Yuqian Zhang, Wenzheng Cheng +7
Surgical phase recognition (SPR) is an integral component of surgical data science, enabling high-level surgical analysis. End-to-end trained neural networks that predict surgical…
DeepDRR -- A Catalyst for Machine Learning in Fluoroscopy-guided Procedures
Mathias Unberath, Jan-Nico Zaech, Sing Chun Lee +4
Machine learning-based approaches outperform competing methods in most disciplines relevant to diagnostic radiology. Interventional radiology, however, has not yet benefited substa…
Generalizing Spatial Transformers to Projective Geometry with Applications to 2D/3D Registration
Cong Gao, Xingtong Liu, Wenhao Gu +4
Differentiable rendering is a technique to connect 3D scenes with corresponding 2D images. Since it is differentiable, processes during image formation can be learned. Previous app…
Explainable AI Enhances Glaucoma Referrals, Yet the Human-AI Team Still Falls Short of the AI Alone
Catalina Gomez, Ruolin Wang, Katharina Breininger +6
Primary care providers are vital for initial triage and referrals to specialty care. In glaucoma, asymptomatic and fast progression can lead to vision loss, necessitating timely re…
Self-supervised Dense 3D Reconstruction from Monocular Endoscopic Video
Xingtong Liu, Ayushi Sinha, Masaru Ishii +3
We present a self-supervised learning-based pipeline for dense 3D reconstruction from full-length monocular endoscopic videos without a priori modeling of anatomy or shading. Our m…
TwinOR: Photorealistic Digital Twins of Dynamic Operating Rooms for Embodied AI Research
Han Zhang, Yiqing Shen, Roger D. Soberanis-Mukul +11
Developing embodied AI for intelligent surgical systems requires safe, controllable environments for continual learning and evaluation. However, safety regulations and operational…
Technical Note: Towards Virtual Monitors for Image Guided Interventions - Real-time Streaming to Optical See-Through Head-Mounted Displays
Long Qian, Mathias Unberath, Kevin Yu +4
Purpose: Image guidance is crucial for the success of many interventions. Images are displayed on designated monitors that cannot be positioned optimally due to sterility and spati…
Pelphix: Surgical Phase Recognition from X-ray Images in Percutaneous Pelvic Fixation
Benjamin D. Killeen, Han Zhang, Jan Mangulabnan +4
Surgical phase recognition (SPR) is a crucial element in the digital transformation of the modern operating theater. While SPR based on video sources is well-established, incorpora…
Constrained Natural Language Action Planning for Resilient Embodied Systems
Grayson Byrd, Corban Rivera, Bethany Kemp +5
Replicating human-level intelligence in the execution of embodied tasks remains challenging due to the unconstrained nature of real-world environments. Novel use of large language…
Pose-dependent weights and Domain Randomization for fully automatic X-ray to CT Registration
Matthias Grimm, Javier Esteban, Mathias Unberath +1
Fully automatic X-ray to CT registration requires a solid initialization to provide an initial alignment within the capture range of existing intensity-based registrations. This wo…
FastSAM-3DSlicer: A 3D-Slicer Extension for 3D Volumetric Segment Anything Model with Uncertainty Quantification
Yiqing Shen, Xinyuan Shao, Blanca Inigo Romillo +2
Accurate segmentation of anatomical structures and pathological regions in medical images is crucial for diagnosis, treatment planning, and disease monitoring. While the Segment An…
SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation
Sampath Rapuri, Lalithkumar Seenivasan, Dominik Schneider +9
A surgical world model capable of generating realistic surgical action videos with precise control over tool-tissue interactions can address fundamental challenges in surgical AI a…
SAGE: SLAM with Appearance and Geometry Prior for Endoscopy
Xingtong Liu, Zhaoshuo Li, Masaru Ishii +3
In endoscopy, many applications (e.g., surgical navigation) would benefit from a real-time method that can simultaneously track the endoscope and reconstruct the dense 3D geometry…
Benchmark of Segmentation Techniques for Pelvic Fracture in CT and X-ray: Summary of the PENGWIN 2024 Challenge
Yudi Sang, Yanzhen Liu, Sutuke Yibulayimu +33
The segmentation of pelvic fracture fragments in CT and X-ray images is crucial for trauma diagnosis, surgical planning, and intraoperative guidance. However, accurately and effici…
Causality-Driven Audits of Model Robustness
Nathan Drenkow, William Paul, Chris Ribaudo +1
Robustness audits of deep neural networks (DNN) provide a means to uncover model sensitivities to the challenging real-world imaging conditions that significantly degrade DNN perfo…
Fast Reasoning Segmentation for Images and Videos
Yiqing Shen, Mathias Unberath
Reasoning segmentation enables open-set object segmentation via implicit text queries, therefore serving as a foundation for embodied agents that should operate autonomously in rea…
Closing the Calibration Loop: An Inside-out-tracking Paradigm for Augmented Reality in Orthopedic Surgery
Jonas Hajek, Mathias Unberath, Javad Fotouhi +6
In percutaneous orthopedic interventions the surgeon attempts to reduce and fixate fractures in bony structures. The complexity of these interventions arises when the surgeon perfo…
Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins
Yiqing Shen, Chenjia Li, Bohan Liu +3
Analyzing operating room (OR) workflows to derive quantitative insights into OR efficiency is important for hospitals to maximize patient care and financial sustainability. Prior w…
On the Sins of Image Synthesis Loss for Self-supervised Depth Estimation
Zhaoshuo Li, Nathan Drenkow, Hao Ding +5
Scene depth estimation from stereo and monocular imagery is critical for extracting 3D information for downstream tasks such as scene understanding. Recently, learning-based method…
Leveraging Vision and Kinematics Data to Improve Realism of Biomechanic Soft-tissue Simulation for Robotic Surgery
Jie Ying Wu, Peter Kazanzides, Mathias Unberath
Purpose Surgical simulations play an increasingly important role in surgeon education and developing algorithms that enable robots to perform surgical subtasks. To model anatomy, F…
E-DSSR: Efficient Dynamic Surgical Scene Reconstruction with Transformer-based Stereoscopic Depth Perception
Yonghao Long, Zhaoshuo Li, Chi Hang Yee +4
Reconstructing the scene of robotic surgery from the stereo endoscopic video is an important and promising topic in surgical data science, which potentially supports many applicati…
The Impact of Machine Learning on 2D/3D Registration for Image-guided Interventions: A Systematic Review and Perspective
Mathias Unberath, Cong Gao, Yicheng Hu +4
Image-based navigation is widely considered the next frontier of minimally invasive surgery. It is believed that image-based navigation will increase the access to reproducible, sa…
A Quantitative Evaluation of Dense 3D Reconstruction of Sinus Anatomy from Monocular Endoscopic Video
Jan Emily Mangulabnan, Roger D. Soberanis-Mukul, Timo Teufel +9
Generating accurate 3D reconstructions from endoscopic video is a promising avenue for longitudinal radiation-free analysis of sinus anatomy and surgical outcomes. Several methods…
Seminar Innovation Management - Winter Term 2017
Gerd Häusler, Aleksandra Milczarek, Markus Schreiter +14
This document contains the results obtained by the Innovation Management Seminar in winter term 2017. In total 11 ideas have been developed by the team. In the document all 11 idea…
Intelligent Control of Robotic X-ray Devices using a Language-promptable Digital Twin
Benjamin D. Killeen, Anushri Suresh, Catalina Gomez +3
Natural language offers a convenient, flexible interface for controlling robotic C-arm X-ray systems, making advanced functionality and controls accessible. However, enabling langu…
Neuralangelo: High-Fidelity Neural Surface Reconstruction
Zhaoshuo Li, Thomas Müller, Alex Evans +4
Neural surface reconstruction has been shown to be powerful for recovering dense 3D surfaces via image-based neural rendering. However, current methods struggle to recover detailed…
Pediatric Otoscopy Video Screening with Shift Contrastive Anomaly Detection
Weiyao Wang, Aniruddha Tamhane, Christine Santos +4
Ear related concerns and symptoms represents the leading indication for seeking pediatric healthcare attention. Despite the high incidence of such encounters, the diagnostic proces…
An Intrinsically Explainable Approach to Detecting Vertebral Compression Fractures in CT Scans via Neurosymbolic Modeling
Blanca Inigo, Yiqing Shen, Benjamin D. Killeen +4
Vertebral compression fractures (VCFs) are a common and potentially serious consequence of osteoporosis. Yet, they often remain undiagnosed. Opportunistic screening, which involves…
Neural Finite-State Machines for Surgical Phase Recognition
Hao Ding, Zhongpai Gao, Benjamin Planche +7
Surgical phase recognition (SPR) is crucial for applications in workflow optimization, performance evaluation, and real-time intervention guidance. However, current deep learning m…
Fast and Automatic Periacetabular Osteotomy Fragment Pose Estimation Using Intraoperatively Implanted Fiducials and Single-View Fluoroscopy
Robert Grupp, Ryan Murphy, Rachel Hegeman +6
Accurate and consistent mental interpretation of fluoroscopy to determine the position and orientation of acetabular bone fragments in 3D space is difficult. We propose a computer…
Task-based Generation of Optimized Projection Sets using Differentiable Ranking
Linda-Sophie Schneider, Mareike Thies, Christopher Syben +3
We present a method for selecting valuable projections in computed tomography (CT) scans to enhance image reconstruction and diagnosis. The approach integrates two important factor…
Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA
Yiqing Shen, Han Zhang, Mathias Unberath
Surgical video question answering requires multi-step reasoning across semantic, spatial, and temporal dimensions. Existing methods architecturally compress videos into discrete to…
Exploring Partial Intrinsic and Extrinsic Symmetry in 3D Medical Imaging
Javad Fotouhi, Giacomo Taylor, Mathias Unberath +5
We present a novel methodology to detect imperfect bilateral symmetry in CT of human anatomy. In this paper, the structurally symmetric nature of the pelvic bone is explored and is…
A Biomechanical Study on the Use of Curved Drilling Technique for Treatment of Osteonecrosis of Femoral Head
Mahsan Bakhtiarinejad, Farshid Alambeigi, Alireza Chamani +3
Osteonecrosis occurs due to the loss of blood supply to the bone, leading to spontaneous death of the trabecular bone. Delayed treatment of the involved patients results in collaps…
Temporally Consistent Online Depth Estimation in Dynamic Scenes
Zhaoshuo Li, Wei Ye, Dilin Wang +4
Temporally consistent depth estimation is crucial for online applications such as augmented reality. While stereo depth estimation has received substantial attention as a promising…
RVTBench: A Benchmark for Visual Reasoning Tasks
Yiqing Shen, Chenjia Li, Chenxiao Fan +1
Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the…
2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Models
Sampath Rapuri, Jeremy Ko, Benjamin D. Killeen +2
The ability to synthesize realistic X-ray images has catalyzed the development of AI models for X-ray image-guided procedures, which otherwise suffer from a lack of available annot…
A County-level Dataset for Informing the United States' Response to COVID-19
Benjamin D. Killeen, Jie Ying Wu, Kinjal Shah +8
As the coronavirus disease 2019 (COVID-19) continues to be a global pandemic, policy makers have enacted and reversed non-pharmaceutical interventions with various levels of restri…
TAToo: Vision-based Joint Tracking of Anatomy and Tool for Skull-base Surgery
Zhaoshuo Li, Hongchao Shu, Ruixing Liang +5
Purpose: Tracking the 3D motion of the surgical tool and the patient anatomy is a fundamental requirement for computer-assisted skull-base surgery. The estimated motion can be used…
Virtual Reality for Synergistic Surgical Training and Data Generation
Adnan Munawar, Zhaoshuo Li, Punit Kunjam +7
Surgical simulators not only allow planning and training of complex procedures, but also offer the ability to generate structured data for algorithm development, which may be appli…
CAI4CAI: The Rise of Contextual Artificial Intelligence in Computer Assisted Interventions
Tom Vercauteren, Mathias Unberath, Nicolas Padoy +1
Data-driven computational approaches have evolved to enable extraction of information from medical images with a reliability, accuracy and speed which is already transforming their…
Learning to Avoid Poor Images: Towards Task-aware C-arm Cone-beam CT Trajectories
Jan-Nico Zaech, Cong Gao, Bastian Bier +4
Metal artifacts in computed tomography (CT) arise from a mismatch between physics of image formation and idealized assumptions during tomographic reconstruction. These artifacts ar…
Data AUDIT: Identifying Attribute Utility- and Detectability-Induced Bias in Task Models
Mitchell Pavlak, Nathan Drenkow, Nicholas Petrick +2
To safely deploy deep learning-based computer vision models for computer-aided detection and diagnosis, we must ensure that they are robust and reliable. Towards that goal, algorit…
CaRTS: Causality-driven Robot Tool Segmentation from Vision and Kinematics Data
Hao Ding, Jintan Zhang, Peter Kazanzides +2
Vision-based segmentation of the robotic tool during robot-assisted surgery enables downstream applications, such as augmented reality feedback, while allowing for inaccuracies in…
Automatic Annotation of Hip Anatomy in Fluoroscopy for Robust and Efficient 2D/3D Registration
Robert Grupp, Mathias Unberath, Cong Gao +7
Fluoroscopy is the standard imaging modality used to guide hip surgery and is therefore a natural sensor for computer-assisted navigation. In order to efficiently solve the complex…
A Learning-based Method for Online Adjustment of C-arm Cone-Beam CT Source Trajectories for Artifact Avoidance
Mareike Thies, Jan-Nico Zäch, Cong Gao +4
During spinal fusion surgery, screws are placed close to critical nerves suggesting the need for highly accurate screw placement. Verifying screw placement on high-quality tomograp…
Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins
Yiqing Shen, Hao Ding, Mathias Unberath
Text-to-video retrieval in operating rooms (OR) is an enabling technology for OR safety, as it allows stakeholders to retrieve and inspect recordings of specific events. However, b…
From Perspective X-ray Imaging to Parallax-Robust Orthographic Stitching
Javad Fotouhi, Xingtong Liu, Mehran Armand +2
Stitching images acquired under perspective projective geometry is a relevant topic in computer vision with multiple applications ranging from smartphone panoramas to the construct…
Did you just see that? Arbitrary view synthesis for egocentric replay of operating room workflows from ambient sensors
Han Zhang, Lalithkumar Seenivasan, Jose L. Porras +9
Observing surgical practice has historically relied on fixed vantage points or recollections, leaving the egocentric visual perspectives that guide clinical decisions undocumented.…
An Interpretable Algorithm for Uveal Melanoma Subtyping from Whole Slide Cytology Images
Haomin Chen, T. Y. Alvin Liu, Catalina Gomez +2
Algorithmic decision support is rapidly becoming a staple of personalized medicine, especially for high-stakes recommendations in which access to certain information can drasticall…
Automated Artifact Detection in Ultra-widefield Fundus Photography of Patients with Sickle Cell Disease
Anqi Feng, Dimitri Johnson, Grace R. Reilly +5
Importance: Ultra-widefield fundus photography (UWF-FP) has shown utility in sickle cell retinopathy screening; however, image artifact may diminish quality and gradeability of ima…
Reasoning Segmentation for Images and Videos: A Survey
Yiqing Shen, Chenjia Li, Fei Xiong +4
Reasoning Segmentation (RS) aims to delineate objects based on implicit text queries, the interpretation of which requires reasoning and knowledge integration. Unlike the tradition…
Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos
Chenyan Jing, Hao Ding, Lalithkumar Seenivasan +2
Vision foundation models such as SAM 3 can provide transferable object-level structure across diverse surgical video conditions, but segmentation outputs do not explicitly encode t…
Towards Controllable Video Synthesis of Routine and Rare OR Events
Dominik Schneider, Lalithkumar Seenivasan, Sampath Rapuri +8
Purpose: Curating large-scale datasets of operating room (OR) workflow, encompassing rare, safety-critical, or atypical events, remains operationally and ethically challenging. Thi…
FluoroSAM: A Language-promptable Foundation Model for Flexible X-ray Image Segmentation
Benjamin D. Killeen, Liam J. Wang, Blanca Inigo +5
Language promptable X-ray image segmentation would enable greater flexibility for human-in-the-loop workflows in diagnostic and interventional precision medicine. Prior efforts hav…
From Generalization to Precision: Exploring SAM for Tool Segmentation in Surgical Environments
Kanyifeechukwu J. Oguine, Roger D. Soberanis-Mukul, Nathan Drenkow +1
Purpose: Accurate tool segmentation is essential in computer-aided procedures. However, this task conveys challenges due to artifacts' presence and the limited training data in med…
Explainable AI for Automated User-specific Feedback in Surgical Skill Acquisition
Catalina Gomez, Lalithkumar Seenivasan, Xinrui Zou +9
Traditional surgical skill acquisition relies heavily on expert feedback, yet direct access is limited by faculty availability and variability in subjective assessments. While trai…
ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and Execution
Corban Rivera, Grayson Byrd, William Paul +12
Robotic planning and execution in open-world environments is a complex problem due to the vast state spaces and high variability of task embodiment. Recent advances in perception a…
Performance and Non-adversarial Robustness of the Segment Anything Model 2 in Surgical Video Segmentation
Yiqing Shen, Hao Ding, Xinyuan Shao +1
Fully supervised deep learning (DL) models for surgical video segmentation have been shown to struggle with non-adversarial, real-world corruptions of image quality including smoke…
Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation
Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding +5
Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation acr…
Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with Transformers
Zhaoshuo Li, Xingtong Liu, Nathan Drenkow +4
Stereo depth estimation relies on optimal correspondence matching between pixels on epipolar lines in the left and right images to infer depth. In this work, we revisit the problem…
Exploiting Partial Structural Symmetry For Patient-Specific Image Augmentation in Trauma Interventions
Javad Fotouhi, Mathias Unberath, Giacomo Taylor +7
In unilateral pelvic fracture reductions, surgeons attempt to reconstruct the bone fragments such that bilateral symmetry in the bony anatomy is restored. We propose to exploit thi…
BronchOpt : Vision-Based Pose Optimization with Fine-Tuned Foundation Models for Accurate Bronchoscopy Navigation
Hongchao Shu, Roger D. Soberanis-Mukul, Jiru Xu +6
Accurate intra-operative localization of the bronchoscope tip relative to patient anatomy remains challenging due to respiratory motion, anatomical variability, and CT-to-body dive…
Beyond Rigid AI: Towards Natural Human-Machine Symbiosis for Interoperative Surgical Assistance
Lalithkumar Seenivasan, Jiru Xu, Roger D. Soberanis Mukul +6
Emerging surgical data science and robotics solutions, especially those designed to provide assistance in situ, require natural human-machine interfaces to fully unlock their poten…
Reflective-AR Display: An Interaction Methodology for Virtual-Real Alignment in Medical Robotics
Javad Fotouhi, Tianyu Song, Arian Mehrfard +8
Robot-assisted minimally invasive surgery has shown to improve patient outcomes, as well as reduce complications and recovery time for several clinical applications. While increasi…
TransNuSeg: A Lightweight Multi-Task Transformer for Nuclei Segmentation
Zhenqi He, Mathias Unberath, Jing Ke +1
Nuclei appear small in size, yet, in real clinical practice, the global spatial information and correlation of the color or brightness contrast between nuclei and background, have…
SyntheX: Scaling Up Learning-based X-ray Image Analysis Through In Silico Experiments
Cong Gao, Benjamin D. Killeen, Yicheng Hu +4
Artificial intelligence (AI) now enables automated interpretation of medical images for clinical use. However, AI's potential use for interventional images (versus those involved i…
Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models
Hao Ding, Lalithkumar Seenivasan, Hongchao Shu +7
Large language model-based (LLM) agents are emerging as a powerful enabler of robust embodied intelligence due to their capability of planning complex action sequences. Sound plann…