Publications (105)
Generic Primitive Detection in Point Clouds Using Novel Minimal Quadric Fits
Tolga Birdal, Benjamin Busam, Nassir Navab +2
We present a novel and effective method for detecting 3D primitives in cluttered, unorganized point clouds, without axillary segmentation or type specification. We consider the qua…
Alignist: CAD-Informed Orientation Distribution Estimation by Fusing Shape and Correspondences
Shishir Reddy Vutukur, Rasmus Laurvig Haugaard, Junwen Huang +2
Object pose distribution estimation is crucial in robotics for better path planning and handling of symmetric objects. Recent distribution estimation approaches employ contrastive…
CloudAttention: Efficient Multi-Scale Attention Scheme For 3D Point Cloud Learning
Mahdi Saleh, Yige Wang, Nassir Navab +2
Processing 3D data efficiently has always been a challenge. Spatial operations on large-scale point clouds, stored as sparse data, require extra cost. Attracted by the success of t…
FLex: Joint Pose and Dynamic Radiance Fields Optimization for Stereo Endoscopic Videos
Florian Philipp Stilz, Mert Asim Karaoglu, Felix Tristram +3
Reconstruction of endoscopic scenes is an important asset for various medical applications, from post-surgery analysis to educational training. Neural rendering has recently shown…
DemoGrasp: Few-Shot Learning for Robotic Grasping with Human Demonstration
Pengyuan Wang, Fabian Manhardt, Luca Minciullo +4
The ability to successfully grasp objects is crucial in robotics, as it enables several interactive downstream applications. To this end, most approaches either compute the full 6D…
RSV: Robotic Sonography for Thyroid Volumetry
John Zielke, Christine Eilers, Benjamin Busam +3
In nuclear medicine, radioiodine therapy is prescribed to treat diseases like hyperthyroidism. The calculation of the prescribed dose depends, amongst other factors, on the thyroid…
Multi-Modal Dataset Acquisition for Photometrically Challenging Object
HyunJun Jung, Patrick Ruhkamp, Nassir Navab +1
This paper addresses the limitations of current datasets for 3D vision tasks in terms of accuracy, size, realism, and suitable imaging modalities for photometrically challenging ob…
Graphite: GRAPH-Induced feaTure Extraction for Point Cloud Registration
Mahdi Saleh, Shervin Dehghani, Benjamin Busam +2
3D Point clouds are a rich source of information that enjoy growing popularity in the vision community. However, due to the sparsity of their representation, learning models based…
CCD-3DR: Consistent Conditioning in Diffusion for Single-Image 3D Reconstruction
Yan Di, Chenyangguang Zhang, Pengyuan Wang +6
In this paper, we present a novel shape reconstruction method leveraging diffusion model to generate 3D sparse point cloud for the object captured in a single RGB image. Recent met…
OperA: Attention-Regularized Transformers for Surgical Phase Recognition
Tobias Czempiel, Magdalini Paschali, Daniel Ostler +3
In this paper we introduce OperA, a transformer-based model that accurately predicts surgical phases from long video sequences. A novel attention regularization loss encourages the…
Rotation-Invariant Transformer for Point Cloud Matching
Hao Yu, Zheng Qin, Ji Hou +4
The intrinsic rotation invariance lies at the core of matching point clouds with handcrafted descriptors. However, it is widely despised by recent deep matchers that obtain the rot…
UnReflectAnything: RGB-Only Highlight Removal by Rendering Synthetic Specular Supervision
Alberto Rota, Mert Kiray, Mert Asim Karaoglu +4
Specular highlights distort appearance, obscure texture, and hinder geometric reasoning in both natural and surgical imagery. We present UnReflectAnything, an RGB-only framework th…
DA Dataset: Toward Dexterity-Aware Dual-Arm Grasping
Guangyao Zhai, Yu Zheng, Ziwei Xu +6
In this paper, we introduce DA, the first large-scale dual-arm dexterity-aware dataset for the generation of optimal bimanual grasping pairs for arbitrary large objects. The da…
ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
Liming Kuang, Yordanka Velikova, Mahdi Saleh +3
Object pose estimation is a fundamental task in computer vision and robotics, yet most methods require extensive, dataset-specific training. Concurrently, large-scale vision langua…
PhysTalk: Language-driven Real-time Physics in 3D Gaussian Scenes
Luca Collorone, Mert Kiray, Indro Spinelli +2
Realistic visual simulations are omnipresent, yet their creation requires computing time, rendering, and expert animation knowledge. Open-vocabulary visual effects generation from…
MonoGraspNet: 6-DoF Grasping with a Single RGB Image
Guangyao Zhai, Dianye Huang, Shun-Cheng Wu +6
6-DoF robotic grasping is a long-lasting but unsolved problem. Recent methods utilize strong 3D networks to extract geometric grasping representations from depth sensors, demonstra…
Reality's Canvas, Language's Brush: Crafting 3D Avatars from Monocular Video
Yuchen Rao, Eduardo Perez Pellitero, Benjamin Busam +2
Recent advancements in 3D avatar generation excel with multi-view supervision for photorealistic models. However, monocular counterparts lag in quality despite broader applicabilit…
CoFiNet: Reliable Coarse-to-fine Correspondences for Robust Point Cloud Registration
Hao Yu, Fu Li, Mahdi Saleh +2
We study the problem of extracting correspondences between a pair of point clouds for registration. For correspondence retrieval, existing works benefit from matching sparse keypoi…
Know your sensORs -- A Modality Study For Surgical Action Classification
Lennart Bastian, Tobias Czempiel, Christian Heiliger +4
The surgical operating room (OR) presents many opportunities for automation and optimization. Videos from various sources in the OR are becoming increasingly available. The medical…
NeRF-Pose: A First-Reconstruct-Then-Regress Approach for Weakly-supervised 6D Object Pose Estimation
Fu Li, Hao Yu, Ivan Shugurov +3
Pose estimation of 3D objects in monocular images is a fundamental and long-standing problem in computer vision. Existing deep learning approaches for 6D pose estimation typically…
On the Importance of Patient Acceptance for Medical Robotic Imaging
Christine Eilers, Rob van Kemenade, Benjamin Busam +1
Purpose: Mutual acceptance is required for any human-to-human interaction. Therefore, one would assume that this also holds for robot-patient interactions. However, the medical rob…
A Minimalist Approach to Type-Agnostic Detection of Quadrics in Point Clouds
Tolga Birdal, Benjamin Busam, Nassir Navab +2
This paper proposes a segmentation-free, automatic and efficient procedure to detect general geometric quadric forms in point clouds, where clutter and occlusions are inevitable. O…
MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
Ege Ãzsoy, Chantal Pellegrini, Tobias Czempiel +7
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assista…
Generative Data Augmentation for Object Point Cloud Segmentation
Dekai Zhu, Stefan Gavranovic, Flavien Boussuge +2
Data augmentation is widely used to train deep learning models to address data scarcity. However, traditional data augmentation (TDA) typically relies on simple geometric transform…
PromptVFX: Text-Driven Fields for Open-World 3D Gaussian Animation
Mert Kiray, Paul Uhlenbruck, Nassir Navab +1
Visual effects (VFX) are key to immersion in modern films, games, and AR/VR. Creating 3D effects requires specialized expertise and training in 3D animation software and can be tim…
CertainNet: Sampling-free Uncertainty Estimation for Object Detection
Stefano Gasperini, Jan Haug, Mohammad-Ali Nikouei Mahani +4
Estimating the uncertainty of a neural network plays a fundamental role in safety-critical settings. In perception for autonomous driving, measuring the uncertainty means providing…
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
Chantal Pellegrini, Ege Ãzsoy, Benjamin Busam +2
Conversational AI tools that can generate and discuss clinically correct radiology reports for a given medical image have the potential to transform radiology. Such a human-in-the-…
FA-BARF: Frequency Adapted Bundle-Adjusting Neural Radiance Fields
Rui Qian, Chenyangguang Zhang, Yan Di +5
Neural Radiance Fields (NeRF) have exhibited highly effective performance for photorealistic novel view synthesis recently. However, the key limitation it meets is the reliance on…
PhoCaL: A Multi-Modal Dataset for Category-Level Object Pose Estimation with Photometrically Challenging Objects
Pengyuan Wang, HyunJun Jung, Yitong Li +6
Object pose estimation is crucial for robotic applications and augmented reality. Beyond instance level 6D object pose estimation methods, estimating category-level pose and shape…
DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
Seok-Young Kim, Abdelrahman Elskhawy, Taewook Ha +4
We present DeWorldSG, a novel framework that generates spatio-temporally robust 3D Semantic Scene Graphs from RGB-D sequences. Existing methods often struggle to construct reliable…
LiteTracker: Leveraging Temporal Causality for Accurate Low-latency Tissue Tracking
Mert Asim Karaoglu, Wenbo Ji, Ahmed Abbas +3
Tissue tracking plays a critical role in various surgical navigation and extended reality (XR) applications. While current methods trained on large synthetic datasets achieve high…
Location-Free Scene Graph Generation
Ege Ãzsoy, Felix Holm, Mahdi Saleh +4
Scene Graph Generation (SGG) is a visual understanding task, aiming to describe a scene as a graph of entities and their relationships with each other. Existing works rely on locat…
ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose Estimation
Yongzhi Su, Mahdi Saleh, Torben Fetzer +5
Establishing correspondences from image to 3D has been a key task of 6DoF object pose estimation for a long time. To predict pose more accurately, deeply learned dense maps replace…
Explaining the Ambiguity of Object Detection and 6D Pose From Visual Data
Fabian Manhardt, Diego Martin Arroyo, Christian Rupprecht +4
3D object detection and pose estimation from a single image are two inherently ambiguous problems. Oftentimes, objects appear similar from different viewpoints due to shape symmetr…
Node-RF: Learning Generalized Continuous Space-Time Scene Dynamics with Neural ODE-based NeRFs
Hiran Sarkar, Liming Kuang, Yordanka Velikova +1
Predicting scene dynamics from visual observations is challenging. Existing methods capture dynamics only within observed boundaries failing to extrapolate far beyond the training…
CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration
Adil Meric, Lin Geng Foo, Mert Kiray +3
We present CoMoGen, a controllable video generation framework that generates realistic interactive dynamics from a single binary mask sequence conditioned on an input image. CoMoGe…
STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
Feng Xu, Guangyao Zhai, Xin Kong +4
Recent advances in Vision-Language-Action (VLA) models, powered by large language models and reinforcement learning-based fine-tuning, have shown remarkable progress in robotic man…
Camera Pose Filtering with Local Regression Geodesics on the Riemannian Manifold of Dual Quaternions
Benjamin Busam, Tolga Birdal, Nassir Navab
Time-varying, smooth trajectory estimation is of great interest to the vision community for accurate and well behaving 3D systems. In this paper, we propose a novel principal compo…
DynaMoN: Motion-Aware Fast and Robust Camera Localization for Dynamic Neural Radiance Fields
Nicolas Schischka, Hannah Schieber, Mert Asim Karaoglu +6
The accurate reconstruction of dynamic scenes with neural radiance fields is significantly dependent on the estimation of camera poses. Widely used structure-from-motion pipelines…
Wild ToFu: Improving Range and Quality of Indirect Time-of-Flight Depth with RGB Fusion in Challenging Environments
HyunJun Jung, Nikolas Brasch, Ales Leonardis +2
Indirect Time-of-Flight (I-ToF) imaging is a widespread way of depth estimation for mobile devices due to its small size and affordable price. Previous works have mainly focused on…
MatchU: Matching Unseen Objects for 6D Pose Estimation from RGB-D Images
Junwen Huang, Hao Yu, Kuan-Ting Yu +3
Recent learning methods for object pose estimation require resource-intensive training for each individual object instance or category, hampering their scalability in real applicat…
OPA-3D: Occlusion-Aware Pixel-Wise Aggregation for Monocular 3D Object Detection
Yongzhi Su, Yan Di, Fabian Manhardt +5
Despite monocular 3D object detection having recently made a significant leap forward thanks to the use of pre-trained depth estimators for pseudo-LiDAR recovery, such two-stage me…
PRISM-0: A Predicate-Rich Scene Graph Generation Framework for Zero-Shot Open-Vocabulary Tasks
Abdelrahman Elskhawy, Mengze Li, Nassir Navab +1
In Scene Graph Generation (SGG), structured representations are extracted from visual inputs as object nodes and connecting predicates, enabling image-based reasoning for diverse d…
HDD-Net: Hybrid Detector Descriptor with Mutual Interactive Learning
Axel Barroso-Laguna, Yannick Verdie, Benjamin Busam +1
Local feature extraction remains an active research area due to the advances in fields such as SLAM, 3D reconstructions, or AR applications. The success in these applications relie…
Polarimetric Pose Prediction
Daoyi Gao, Yitong Li, Patrick Ruhkamp +6
Light has many properties that vision sensors can passively measure. Colour-band separated wavelength and intensity are arguably the most commonly used for monocular 6D object pose…
Deformable 3D Gaussian Splatting for Animatable Human Avatars
HyunJun Jung, Nikolas Brasch, Jifei Song +5
Recent advances in neural radiance fields enable novel view synthesis of photo-realistic images in dynamic settings, which can be applied to scenarios with human animation. Commonl…
DisguisOR: Holistic Face Anonymization for the Operating Room
Lennart Bastian, Tony Danjun Wang, Tobias Czempiel +2
Purpose: Recent advances in Surgical Data Science (SDS) have contributed to an increase in video recordings from hospital environments. While methods such as surgical workflow reco…
Pose Anything Anywhere:Model-free Object Poses from Arbitrary References
Hongli Xu, Jiaqi Hu, Junwen Huang +5
Estimating the 6D pose of unseen objects is a fundamental yet challenging problem for open-world robotics and embodied perception. Model-based methods are accurate but depend on CA…
Robotic Navigation Autonomy for Subretinal Injection via Intelligent Real-Time Virtual iOCT Volume Slicing
Shervin Dehghani, Michael Sommersperger, Peiyao Zhang +6
In the last decade, various robotic platforms have been introduced that could support delicate retinal surgeries. Concurrently, to provide semantic understanding of the surgical ar…
On the Localization of Ultrasound Image Slices within Point Distribution Models
Lennart Bastian, Vincent Bürgin, Ha Young Kim +4
Thyroid disorders are most commonly diagnosed using high-resolution Ultrasound (US). Longitudinal nodule tracking is a pivotal diagnostic protocol for monitoring changes in patholo…
SteReFo: Efficient Image Refocusing with Stereo Vision
Benjamin Busam, Matthieu Hog, Steven McDonagh +1
Whether to attract viewer attention to a particular object, give the impression of depth or simply reproduce human-like scene perception, shallow depth of field images are used ext…
VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models
Shuhao Kang, Youqi Liao, Peijie Wang +5
Text-to-point-cloud (T2P) localization aims to infer precise spatial positions within 3D point cloud maps from natural language descriptions, reflecting how humans perceive and com…
CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion
Guangyao Zhai, Evin Pınar Ãrnek, Shun-Cheng Wu +4
Controllable scene synthesis aims to create interactive environments for various industrial use cases. Scene graphs provide a highly suitable interface to facilitate these applicat…
Supercharging Thermal Gaussian Splatting with Depth Estimation
Manoj Biswanath, Chenxin Cai, Hannah Schieber +2
Efficient and robust 3D scene representation is crucial in autonomous driving, robotics, and related fields. While RGB images provide valuable content for 3D reconstruction, other…
OSOP: A Multi-Stage One Shot Object Pose Estimation Framework
Ivan Shugurov, Fu Li, Benjamin Busam +1
We present a novel one-shot method for object detection and 6 DoF pose estimation, that does not require training on target objects. At test time, it takes as input a target image…
A Multi-Hypothesis Approach to Color Constancy
Daniel Hernandez-Juarez, Sarah Parisot, Benjamin Busam +3
Contemporary approaches frame the color constancy problem as learning camera specific illuminant mappings. While high accuracy can be achieved on camera specific data, these models…
I Like to Move It: 6D Pose Estimation as an Action Decision Process
Benjamin Busam, Hyun Jun Jung, Nassir Navab
Object pose estimation is an integral part of robot vision and AR. Previous 6D pose retrieval pipelines treat the problem either as a regression task or discretize the pose space t…
3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection
Alexander Lehner, Stefano Gasperini, Alvaro Marcos-Ramiro +5
As 3D object detection on point clouds relies on the geometrical relationships between the points, non-standard object shapes can hinder a method's detection capability. However, i…
UnderOneFacade: Worldwide Facade Semantic Segmentation Benchmark Dataset
Yi Wang, Fan Wang, Prabin Gyawali +9
Globally consistent semantic digital twins require centimeter-accurate and geographically transferable 3D facade segmentation. However, progress in facade parsing is limited by the…
Object Pose Transformer: Unifying Unseen Object Pose Estimation
Weihang Li, Lorenzo Garattoni, Fabien Despinoy +2
Learning model-free object pose estimation for unseen instances remains a fundamental challenge in 3D vision. Existing methods typically fall into two disjoint paradigms: category-…
Ultra-NeRF: Neural Radiance Fields for Ultrasound Imaging
Magdalena Wysocki, Mohammad Farid Azampour, Christine Eilers +3
We present a physics-enhanced implicit neural representation (INR) for ultrasound (US) imaging that learns tissue properties from overlapping US sweeps. Our proposed method leverag…
Dropping the D: RGB-D SLAM Without the Depth Sensor
Mert Kiray, Alican Karaomer, Benjamin Busam
We present DropD-SLAM, a real-time monocular SLAM system that achieves RGB-D-level accuracy without relying on depth sensors. The system replaces active depth input with three pret…
RIGA: Rotation-Invariant and Globally-Aware Descriptors for Point Cloud Registration
Hao Yu, Ji Hou, Zheng Qin +5
Successful point cloud registration relies on accurate correspondences established upon powerful descriptors. However, existing neural descriptors either leverage a rotation-varian…
Markerless Inside-Out Tracking for Interventional Applications
Benjamin Busam, Patrick Ruhkamp, Salvatore Virga +4
Tracking of rotation and translation of medical instruments plays a substantial role in many modern interventions. Traditional external optical tracking systems are often subject t…
CPS++: Improving Class-level 6D Pose and Shape Estimation From Monocular Images With Self-Supervised Learning
Fabian Manhardt, Gu Wang, Benjamin Busam +5
Contemporary monocular 6D pose estimation methods can only cope with a handful of object instances. This naturally hampers possible applications as, for instance, robots seamlessly…
ODeform: Learning Continuous 4D Motion for Shape Deformation with Neural ODEs
Yordanka Velikova, Mahdi Saleh, Liming Kuang +1
Modeling continuous object deformation is important for many computer vision and robotics tasks, such as manipulation and simulation. Existing approaches rely on learning-based met…
On the Importance of Accurate Geometry Data for Dense 3D Vision Tasks
HyunJun Jung, Patrick Ruhkamp, Guangyao Zhai +10
Learning-based methods to solve dense 3D vision problems typically train on 3D sensor data. The respectively used principle of measuring distances provides advantages and drawbacks…
GS4City: Hierarchical Semantic Gaussian Splatting via City-Model Priors
Qilin Zhang, Jinyu Zhu, Olaf Wysocki +2
Recent semantic 3D Gaussian Splatting (3DGS) methods primarily rely on 2D foundation models, often yielding ambiguous boundaries and limited support for structured urban semantics.…
DynaMiTe: A Dynamic Local Motion Model with Temporal Constraints for Robust Real-Time Feature Matching
Patrick Ruhkamp, Ruiqi Gong, Nassir Navab +1
Feature based visual odometry and SLAM methods require accurate and fast correspondence matching between consecutive image frames for precise camera pose estimation in real-time. C…
CroMo: Cross-Modal Learning for Monocular Depth Estimation
Yannick Verdié, Jifei Song, Barnabé Mas +3
Learning-based depth estimation has witnessed recent progress in multiple directions; from self-supervision using monocular video to supervised methods offering highest accuracy. C…
Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
Guangyao Zhai, Yue Zhou, Xinyan Deng +3
Few-shot anomaly detection streamlines and simplifies industrial safety inspection. However, limited samples make accurate differentiation between normal and abnormal features chal…
NeRF-Feat: 6D Object Pose Estimation using Feature Rendering
Shishir Reddy Vutukur, Heike Brock, Benjamin Busam +3
Object Pose Estimation is a crucial component in robotic grasping and augmented reality. Learning based approaches typically require training data from a highly accurate CAD model…
BioDet: Boosting Industrial Object Detection with Image Preprocessing Strategies
Jiaqi Hu, Hongli Xu, Junwen Huang +3
Accurate 6D pose estimation is essential for robotic manipulation in industrial environments. Existing pipelines typically rely on off-the-shelf object detectors followed by croppi…
RIDE: Self-Supervised Learning of Rotation-Equivariant Keypoint Detection and Invariant Description for Endoscopy
Mert Asim Karaoglu, Viktoria Markova, Nassir Navab +2
Unlike in natural images, in endoscopy there is no clear notion of an up-right camera orientation. Endoscopic videos therefore often contain large rotational motions, which require…
Project to Adapt: Domain Adaptation for Depth Completion from Noisy and Sparse Sensor Data
Adrian Lopez-Rodriguez, Benjamin Busam, Krystian Mikolajczyk
Depth completion aims to predict a dense depth map from a sparse depth input. The acquisition of dense ground truth annotations for depth completion settings can be difficult and,…
EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion
Guangyao Zhai, Evin Pınar Ãrnek, Dave Zhenyu Chen +5
We present EchoScene, an interactive and controllable generative model that generates 3D indoor scenes on scene graphs. EchoScene leverages a dual-branch diffusion model that dynam…
BFS-Net: Weakly Supervised Cell Instance Segmentation from Bright-Field Microscopy Z-Stacks
Shervin Dehghani, Benjamin Busam, Nassir Navab +1
Despite its broad availability, volumetric information acquisition from Bright-Field Microscopy (BFM) is inherently difficult due to the projective nature of the acquisition proces…
Disentangling 3D Attributes from a Single 2D Image: Human Pose, Shape and Garment
Xue Hu, Xinghui Li, Benjamin Busam +3
For visual manipulation tasks, we aim to represent image content with semantically meaningful features. However, learning implicit representations from images often lacks interpret…
SABER-6D: Shape Representation Based Implicit Object Pose Estimation
Shishir Reddy Vutukur, Mengkejiergeli Ba, Benjamin Busam +2
In this paper, we propose a novel encoder-decoder architecture, named SABER, to learn the 6D pose of the object in the embedding space by learning shape representation at a given p…
Polarimetric Information for Multi-Modal 6D Pose Estimation of Photometrically Challenging Objects with Limited Data
Patrick Ruhkamp, Daoyi Gao, HyunJun Jung +2
6D pose estimation pipelines that rely on RGB-only or RGB-D data show limitations for photometrically challenging objects with e.g. textureless surfaces, reflections or transparenc…
Bending Graphs: Hierarchical Shape Matching using Gated Optimal Transport
Mahdi Saleh, Shun-Cheng Wu, Luca Cosmo +3
Shape matching has been a long-studied problem for the computer graphics and vision community. The objective is to predict a dense correspondence between meshes that have a certain…
MultiCam: On-the-fly Multi-Camera Pose Estimation Using Spatiotemporal Overlaps of Known Objects
Shiyu Li, Hannah Schieber, Kristoffer Waldow +3
Multi-camera dynamic Augmented Reality (AR) applications require a camera pose estimation to leverage individual information from each camera in one common system. This can be achi…
3D Adversarial Augmentations for Robust Out-of-Domain Predictions
Alexander Lehner, Stefano Gasperini, Alvaro Marcos-Ramiro +4
Since real-world training datasets cannot properly sample the long tail of the underlying data distribution, corner cases and rare out-of-domain samples can severely hinder the per…
3D Segmentation Using Viewpoint-Dependent Spatial Relationships
Ayaka Nanri, Klara Reichard, Mert Kiray +3
Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentation methods do not explicitl…
GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation
Weihang Li, Hongli Xu, Junwen Huang +4
A key challenge in model-free category-level pose estimation is the extraction of contextual object features that generalize across varying instances within a specific category. Re…
S3M: Scalable Statistical Shape Modeling through Unsupervised Correspondences
Lennart Bastian, Alexander Baumann, Emily Hoppe +5
Statistical shape models (SSMs) are an established way to represent the anatomy of a population with various clinically relevant applications. However, they typically require domai…
Zero123-6D: Zero-shot Novel View Synthesis for RGB Category-level 6D Pose Estimation
Francesco Di Felice, Alberto Remus, Stefano Gasperini +5
Estimating the pose of objects through vision is essential to make robotic platforms interact with the environment. Yet, it presents many challenges, often related to the lack of f…
SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation
Yamei Chen, Yan Di, Guangyao Zhai +6
Category-level object pose estimation, aiming to predict the 6D pose and 3D size of objects from known categories, typically struggles with large intra-class shape variation. Exist…
Segmenting Known Objects and Unseen Unknowns without Prior Knowledge
Stefano Gasperini, Alvaro Marcos-Ramiro, Michael Schmidt +3
Panoptic segmentation methods assign a known class to each pixel given in input. Even for state-of-the-art approaches, this inevitably enforces decisions that systematically lead t…
Generative 6D Pose Estimation via Conditional Flow Matching
Amir Hamza, Davide Boscaini, Weihang Li +2
Existing methods for instance-level 6D pose estimation typically rely on neural networks that either directly regress the pose in or estimate it indirectly via loc…
TexPose: Neural Texture Learning for Self-Supervised 6D Object Pose Estimation
Hanzhi Chen, Fabian Manhardt, Nassir Navab +1
In this paper, we introduce neural texture learning for 6D object pose estimation from synthetic data and a few unlabelled real images. Our major contribution is a novel learning s…
R4Dyn: Exploring Radar for Self-Supervised Monocular Depth Estimation of Dynamic Scenes
Stefano Gasperini, Patrick Koch, Vinzenz Dallabetta +3
While self-supervised monocular depth estimation in driving scenarios has achieved comparable performance to supervised approaches, violations of the static world assumption can st…
Attention meets Geometry: Geometry Guided Spatial-Temporal Attention for Consistent Self-Supervised Monocular Depth Estimation
Patrick Ruhkamp, Daoyi Gao, Hanzhi Chen +2
Inferring geometrically consistent dense 3D scenes across a tuple of temporally consecutive images remains challenging for self-supervised monocular depth prediction pipelines. Thi…
BenchSeg: A Large-Scale Dataset and Benchmark for Multi-View Food Video Segmentation
Ahmad AlMughrabi, Guillermo Rivo, Carlos Jiménez-Farfán +6
Food image segmentation is a critical task for dietary analysis, enabling accurate estimation of food volume and nutrients. However, current methods suffer from limited multi-view…
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
Ege Ãzsoy, Arda Mamur, Felix Tristram +4
Operating rooms (ORs) demand precise coordination among surgeons, nurses, and equipment in a fast-paced, occlusion-heavy environment, necessitating advanced perception models to en…
S2P3: Self-Supervised Polarimetric Pose Prediction
Patrick Ruhkamp, Daoyi Gao, Nassir Navab +1
This paper proposes the first self-supervised 6D object pose prediction from multimodal RGB+polarimetric images. The novel training paradigm comprises 1) a physical model to extrac…
MS-CLR: Multi-Skeleton Contrastive Learning for Human Action Recognition
Mert Kiray, Alvaro Ritter, Nassir Navab +1
Contrastive learning has gained significant attention in skeleton-based action recognition for its ability to learn robust representations from unlabeled data. However, existing me…
Is my Depth Ground-Truth Good Enough? HAMMER -- Highly Accurate Multi-Modal Dataset for DEnse 3D Scene Regression
HyunJun Jung, Patrick Ruhkamp, Guangyao Zhai +9
Depth estimation is a core task in 3D computer vision. Recent methods investigate the task of monocular depth trained with various depth sensor modalities. Every sensor has its adv…
ColibriDoc: An Eye-in-Hand Autonomous Trocar Docking System
Shervin Dehghani, Michael Sommersperger, Junjie Yang +6
Retinal surgery is a complex medical procedure that requires exceptional expertise and dexterity. For this purpose, several robotic platforms are currently being developed to enabl…
RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation
Junwen Huang, Shishir Reddy Vutukur, Peter KT Yu +3
Typical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the…