papers

Publications (105)

cs.CV2019

Generic Primitive Detection in Point Clouds Using Novel Minimal Quadric Fits

Tolga Birdal, Benjamin Busam, Nassir Navab +2

We present a novel and effective method for detecting 3D primitives in cluttered, unorganized point clouds, without axillary segmentation or type specification. We consider the qua…

cs.CV2024

Alignist: CAD-Informed Orientation Distribution Estimation by Fusing Shape and Correspondences

Shishir Reddy Vutukur, Rasmus Laurvig Haugaard, Junwen Huang +2

Object pose distribution estimation is crucial in robotics for better path planning and handling of symmetric objects. Recent distribution estimation approaches employ contrastive…

cs.CV2022

CloudAttention: Efficient Multi-Scale Attention Scheme For 3D Point Cloud Learning

Mahdi Saleh, Yige Wang, Nassir Navab +2

Processing 3D data efficiently has always been a challenge. Spatial operations on large-scale point clouds, stored as sparse data, require extra cost. Attracted by the success of t…

cs.CV2024

FLex: Joint Pose and Dynamic Radiance Fields Optimization for Stereo Endoscopic Videos

Florian Philipp Stilz, Mert Asim Karaoglu, Felix Tristram +3

Reconstruction of endoscopic scenes is an important asset for various medical applications, from post-surgery analysis to educational training. Neural rendering has recently shown…

cs.RO2021

DemoGrasp: Few-Shot Learning for Robotic Grasping with Human Demonstration

Pengyuan Wang, Fabian Manhardt, Luca Minciullo +4

The ability to successfully grasp objects is crucial in robotics, as it enables several interactive downstream applications. To this end, most approaches either compute the full 6D…

cs.RO2021

RSV: Robotic Sonography for Thyroid Volumetry

John Zielke, Christine Eilers, Benjamin Busam +3

In nuclear medicine, radioiodine therapy is prescribed to treat diseases like hyperthyroidism. The calculation of the prescribed dose depends, amongst other factors, on the thyroid…

cs.CV2023

Multi-Modal Dataset Acquisition for Photometrically Challenging Object

HyunJun Jung, Patrick Ruhkamp, Nassir Navab +1

This paper addresses the limitations of current datasets for 3D vision tasks in terms of accuracy, size, realism, and suitable imaging modalities for photometrically challenging ob…

cs.CV2020

Graphite: GRAPH-Induced feaTure Extraction for Point Cloud Registration

Mahdi Saleh, Shervin Dehghani, Benjamin Busam +2

3D Point clouds are a rich source of information that enjoy growing popularity in the vision community. However, due to the sparsity of their representation, learning models based…

cs.CV2023

CCD-3DR: Consistent Conditioning in Diffusion for Single-Image 3D Reconstruction

Yan Di, Chenyangguang Zhang, Pengyuan Wang +6

In this paper, we present a novel shape reconstruction method leveraging diffusion model to generate 3D sparse point cloud for the object captured in a single RGB image. Recent met…

cs.CV2021

OperA: Attention-Regularized Transformers for Surgical Phase Recognition

Tobias Czempiel, Magdalini Paschali, Daniel Ostler +3

In this paper we introduce OperA, a transformer-based model that accurately predicts surgical phases from long video sequences. A novel attention regularization loss encourages the…

cs.CV2024

Rotation-Invariant Transformer for Point Cloud Matching

Hao Yu, Zheng Qin, Ji Hou +4

The intrinsic rotation invariance lies at the core of matching point clouds with handcrafted descriptors. However, it is widely despised by recent deep matchers that obtain the rot…

cs.CV2025

UnReflectAnything: RGB-Only Highlight Removal by Rendering Synthetic Specular Supervision

Alberto Rota, Mert Kiray, Mert Asim Karaoglu +4

Specular highlights distort appearance, obscure texture, and hinder geometric reasoning in both natural and surgical imagery. We present UnReflectAnything, an RGB-only framework th…

cs.RO2022

DA Dataset: Toward Dexterity-Aware Dual-Arm Grasping

Guangyao Zhai, Yu Zheng, Ziwei Xu +6

In this paper, we introduce DA, the first large-scale dual-arm dexterity-aware dataset for the generation of optimal bimanual grasping pairs for arbitrary large objects. The da…

cs.CV2026

ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors

Liming Kuang, Yordanka Velikova, Mahdi Saleh +3

Object pose estimation is a fundamental task in computer vision and robotics, yet most methods require extensive, dataset-specific training. Concurrently, large-scale vision langua…

cs.GR2025

PhysTalk: Language-driven Real-time Physics in 3D Gaussian Scenes

Luca Collorone, Mert Kiray, Indro Spinelli +2

Realistic visual simulations are omnipresent, yet their creation requires computing time, rendering, and expert animation knowledge. Open-vocabulary visual effects generation from…

cs.RO2023

MonoGraspNet: 6-DoF Grasping with a Single RGB Image

Guangyao Zhai, Dianye Huang, Shun-Cheng Wu +6

6-DoF robotic grasping is a long-lasting but unsolved problem. Recent methods utilize strong 3D networks to extract geometric grasping representations from depth sensors, demonstra…

cs.CV2024

Reality's Canvas, Language's Brush: Crafting 3D Avatars from Monocular Video

Yuchen Rao, Eduardo Perez Pellitero, Benjamin Busam +2

Recent advancements in 3D avatar generation excel with multi-view supervision for photorealistic models. However, monocular counterparts lag in quality despite broader applicabilit…

cs.CV2021

CoFiNet: Reliable Coarse-to-fine Correspondences for Robust Point Cloud Registration

Hao Yu, Fu Li, Mahdi Saleh +2

We study the problem of extracting correspondences between a pair of point clouds for registration. For correspondence retrieval, existing works benefit from matching sparse keypoi…

cs.CV2022

Know your sensORs -- A Modality Study For Surgical Action Classification

Lennart Bastian, Tobias Czempiel, Christian Heiliger +4

The surgical operating room (OR) presents many opportunities for automation and optimization. Videos from various sources in the OR are becoming increasingly available. The medical…

cs.CV2023

NeRF-Pose: A First-Reconstruct-Then-Regress Approach for Weakly-supervised 6D Object Pose Estimation

Fu Li, Hao Yu, Ivan Shugurov +3

Pose estimation of 3D objects in monocular images is a fundamental and long-standing problem in computer vision. Existing deep learning approaches for 6D pose estimation typically…

cs.RO2023

On the Importance of Patient Acceptance for Medical Robotic Imaging

Christine Eilers, Rob van Kemenade, Benjamin Busam +1

Purpose: Mutual acceptance is required for any human-to-human interaction. Therefore, one would assume that this also holds for robot-patient interactions. However, the medical rob…

cs.CV2018

A Minimalist Approach to Type-Agnostic Detection of Quadrics in Point Clouds

Tolga Birdal, Benjamin Busam, Nassir Navab +2

This paper proposes a segmentation-free, automatic and efficient procedure to detect general geometric quadric forms in point clouds, where clutter and occlusions are inevitable. O…

cs.CV2025

MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments

Ege Özsoy, Chantal Pellegrini, Tobias Czempiel +7

Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assista…

cs.CV2025

Generative Data Augmentation for Object Point Cloud Segmentation

Dekai Zhu, Stefan Gavranovic, Flavien Boussuge +2

Data augmentation is widely used to train deep learning models to address data scarcity. However, traditional data augmentation (TDA) typically relies on simple geometric transform…

cs.GR2026

PromptVFX: Text-Driven Fields for Open-World 3D Gaussian Animation

Mert Kiray, Paul Uhlenbruck, Nassir Navab +1

Visual effects (VFX) are key to immersion in modern films, games, and AR/VR. Creating 3D effects requires specialized expertise and training in 3D animation software and can be tim…

cs.CV2021

CertainNet: Sampling-free Uncertainty Estimation for Object Detection

Stefano Gasperini, Jan Haug, Mohammad-Ali Nikouei Mahani +4

Estimating the uncertainty of a neural network plays a fundamental role in safety-critical settings. In perception for autonomous driving, measuring the uncertainty means providing…

cs.CV2025

RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance

Chantal Pellegrini, Ege Özsoy, Benjamin Busam +2

Conversational AI tools that can generate and discuss clinically correct radiology reports for a given medical image have the potential to transform radiology. Such a human-in-the-…

cs.CV2025

FA-BARF: Frequency Adapted Bundle-Adjusting Neural Radiance Fields

Rui Qian, Chenyangguang Zhang, Yan Di +5

Neural Radiance Fields (NeRF) have exhibited highly effective performance for photorealistic novel view synthesis recently. However, the key limitation it meets is the reliance on…

cs.CV2022

PhoCaL: A Multi-Modal Dataset for Category-Level Object Pose Estimation with Photometrically Challenging Objects

Pengyuan Wang, HyunJun Jung, Yitong Li +6

Object pose estimation is crucial for robotic applications and augmented reality. Beyond instance level 6D object pose estimation methods, estimating category-level pose and shape…

cs.CV2026

DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors

Seok-Young Kim, Abdelrahman Elskhawy, Taewook Ha +4

We present DeWorldSG, a novel framework that generates spatio-temporally robust 3D Semantic Scene Graphs from RGB-D sequences. Existing methods often struggle to construct reliable…

cs.CV2025

LiteTracker: Leveraging Temporal Causality for Accurate Low-latency Tissue Tracking

Mert Asim Karaoglu, Wenbo Ji, Ahmed Abbas +3

Tissue tracking plays a critical role in various surgical navigation and extended reality (XR) applications. While current methods trained on large synthetic datasets achieve high…

cs.CV2025

Location-Free Scene Graph Generation

Ege Özsoy, Felix Holm, Mahdi Saleh +4

Scene Graph Generation (SGG) is a visual understanding task, aiming to describe a scene as a graph of entities and their relationships with each other. Existing works rely on locat…

cs.CV2022

ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose Estimation

Yongzhi Su, Mahdi Saleh, Torben Fetzer +5

Establishing correspondences from image to 3D has been a key task of 6DoF object pose estimation for a long time. To predict pose more accurately, deeply learned dense maps replace…

cs.CV2019

Explaining the Ambiguity of Object Detection and 6D Pose From Visual Data

Fabian Manhardt, Diego Martin Arroyo, Christian Rupprecht +4

3D object detection and pose estimation from a single image are two inherently ambiguous problems. Oftentimes, objects appear similar from different viewpoints due to shape symmetr…

cs.CV2026

Node-RF: Learning Generalized Continuous Space-Time Scene Dynamics with Neural ODE-based NeRFs

Hiran Sarkar, Liming Kuang, Yordanka Velikova +1

Predicting scene dynamics from visual observations is challenging. Existing methods capture dynamics only within observed boundaries failing to extrapolate far beyond the training…

cs.CV2026

CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration

Adil Meric, Lin Geng Foo, Mert Kiray +3

We present CoMoGen, a controllable video generation framework that generates realistic interactive dynamics from a single binary mask sequence conditioned on an input image. CoMoGe…

cs.RO2025

STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models

Feng Xu, Guangyao Zhai, Xin Kong +4

Recent advances in Vision-Language-Action (VLA) models, powered by large language models and reinforcement learning-based fine-tuning, have shown remarkable progress in robotic man…

cs.CV2017

Camera Pose Filtering with Local Regression Geodesics on the Riemannian Manifold of Dual Quaternions

Benjamin Busam, Tolga Birdal, Nassir Navab

Time-varying, smooth trajectory estimation is of great interest to the vision community for accurate and well behaving 3D systems. In this paper, we propose a novel principal compo…

cs.CV2024

DynaMoN: Motion-Aware Fast and Robust Camera Localization for Dynamic Neural Radiance Fields

Nicolas Schischka, Hannah Schieber, Mert Asim Karaoglu +6

The accurate reconstruction of dynamic scenes with neural radiance fields is significantly dependent on the estimation of camera poses. Widely used structure-from-motion pipelines…

cs.CV2022

Wild ToFu: Improving Range and Quality of Indirect Time-of-Flight Depth with RGB Fusion in Challenging Environments

HyunJun Jung, Nikolas Brasch, Ales Leonardis +2

Indirect Time-of-Flight (I-ToF) imaging is a widespread way of depth estimation for mobile devices due to its small size and affordable price. Previous works have mainly focused on…

cs.CV2024

MatchU: Matching Unseen Objects for 6D Pose Estimation from RGB-D Images

Junwen Huang, Hao Yu, Kuan-Ting Yu +3

Recent learning methods for object pose estimation require resource-intensive training for each individual object instance or category, hampering their scalability in real applicat…

cs.CV2022

OPA-3D: Occlusion-Aware Pixel-Wise Aggregation for Monocular 3D Object Detection

Yongzhi Su, Yan Di, Fabian Manhardt +5

Despite monocular 3D object detection having recently made a significant leap forward thanks to the use of pre-trained depth estimators for pseudo-LiDAR recovery, such two-stage me…

cs.CV2025

PRISM-0: A Predicate-Rich Scene Graph Generation Framework for Zero-Shot Open-Vocabulary Tasks

Abdelrahman Elskhawy, Mengze Li, Nassir Navab +1

In Scene Graph Generation (SGG), structured representations are extracted from visual inputs as object nodes and connecting predicates, enabling image-based reasoning for diverse d…

cs.CV2020

HDD-Net: Hybrid Detector Descriptor with Mutual Interactive Learning

Axel Barroso-Laguna, Yannick Verdie, Benjamin Busam +1

Local feature extraction remains an active research area due to the advances in fields such as SLAM, 3D reconstructions, or AR applications. The success in these applications relie…

cs.CV2022

Polarimetric Pose Prediction

Daoyi Gao, Yitong Li, Patrick Ruhkamp +6

Light has many properties that vision sensors can passively measure. Colour-band separated wavelength and intensity are arguably the most commonly used for monocular 6D object pose…

cs.CV2023

Deformable 3D Gaussian Splatting for Animatable Human Avatars

HyunJun Jung, Nikolas Brasch, Jifei Song +5

Recent advances in neural radiance fields enable novel view synthesis of photo-realistic images in dynamic settings, which can be applied to scenarios with human animation. Commonl…

cs.CV2023

DisguisOR: Holistic Face Anonymization for the Operating Room

Lennart Bastian, Tony Danjun Wang, Tobias Czempiel +2

Purpose: Recent advances in Surgical Data Science (SDS) have contributed to an increase in video recordings from hospital environments. While methods such as surgical workflow reco…

cs.CV2026

Pose Anything Anywhere:Model-free Object Poses from Arbitrary References

Hongli Xu, Jiaqi Hu, Junwen Huang +5

Estimating the 6D pose of unseen objects is a fundamental yet challenging problem for open-world robotics and embodied perception. Model-based methods are accurate but depend on CA…

cs.RO2023

Robotic Navigation Autonomy for Subretinal Injection via Intelligent Real-Time Virtual iOCT Volume Slicing

Shervin Dehghani, Michael Sommersperger, Peiyao Zhang +6

In the last decade, various robotic platforms have been introduced that could support delicate retinal surgeries. Concurrently, to provide semantic understanding of the surgical ar…

eess.IV2023

On the Localization of Ultrasound Image Slices within Point Distribution Models

Lennart Bastian, Vincent Bürgin, Ha Young Kim +4

Thyroid disorders are most commonly diagnosed using high-resolution Ultrasound (US). Longitudinal nodule tracking is a pivotal diagnostic protocol for monitoring changes in patholo…

cs.CV2019

SteReFo: Efficient Image Refocusing with Stereo Vision

Benjamin Busam, Matthieu Hog, Steven McDonagh +1

Whether to attract viewer attention to a particular object, give the impression of depth or simply reproduce human-like scene perception, shallow depth of field images are used ext…

cs.CV2026

VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models

Shuhao Kang, Youqi Liao, Peijie Wang +5

Text-to-point-cloud (T2P) localization aims to infer precise spatial positions within 3D point cloud maps from natural language descriptions, reflecting how humans perceive and com…

cs.CV2023

CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion

Guangyao Zhai, Evin Pınar Örnek, Shun-Cheng Wu +4

Controllable scene synthesis aims to create interactive environments for various industrial use cases. Scene graphs provide a highly suitable interface to facilitate these applicat…

cs.CV2026

Supercharging Thermal Gaussian Splatting with Depth Estimation

Manoj Biswanath, Chenxin Cai, Hannah Schieber +2

Efficient and robust 3D scene representation is crucial in autonomous driving, robotics, and related fields. While RGB images provide valuable content for 3D reconstruction, other…

cs.CV2022

OSOP: A Multi-Stage One Shot Object Pose Estimation Framework

Ivan Shugurov, Fu Li, Benjamin Busam +1

We present a novel one-shot method for object detection and 6 DoF pose estimation, that does not require training on target objects. At test time, it takes as input a target image…

cs.CV2020

A Multi-Hypothesis Approach to Color Constancy

Daniel Hernandez-Juarez, Sarah Parisot, Benjamin Busam +3

Contemporary approaches frame the color constancy problem as learning camera specific illuminant mappings. While high accuracy can be achieved on camera specific data, these models…

cs.CV2020

I Like to Move It: 6D Pose Estimation as an Action Decision Process

Benjamin Busam, Hyun Jun Jung, Nassir Navab

Object pose estimation is an integral part of robot vision and AR. Previous 6D pose retrieval pipelines treat the problem either as a regression task or discretize the pose space t…

cs.CV2022

3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection

Alexander Lehner, Stefano Gasperini, Alvaro Marcos-Ramiro +5

As 3D object detection on point clouds relies on the geometrical relationships between the points, non-standard object shapes can hinder a method's detection capability. However, i…

cs.CV2026

UnderOneFacade: Worldwide Facade Semantic Segmentation Benchmark Dataset

Yi Wang, Fan Wang, Prabin Gyawali +9

Globally consistent semantic digital twins require centimeter-accurate and geographically transferable 3D facade segmentation. However, progress in facade parsing is limited by the…

cs.CV2026

Object Pose Transformer: Unifying Unseen Object Pose Estimation

Weihang Li, Lorenzo Garattoni, Fabien Despinoy +2

Learning model-free object pose estimation for unseen instances remains a fundamental challenge in 3D vision. Existing methods typically fall into two disjoint paradigms: category-…

eess.IV2023

Ultra-NeRF: Neural Radiance Fields for Ultrasound Imaging

Magdalena Wysocki, Mohammad Farid Azampour, Christine Eilers +3

We present a physics-enhanced implicit neural representation (INR) for ultrasound (US) imaging that learns tissue properties from overlapping US sweeps. Our proposed method leverag…

cs.CV2025

Dropping the D: RGB-D SLAM Without the Depth Sensor

Mert Kiray, Alican Karaomer, Benjamin Busam

We present DropD-SLAM, a real-time monocular SLAM system that achieves RGB-D-level accuracy without relying on depth sensors. The system replaces active depth input with three pret…

cs.CV2022

RIGA: Rotation-Invariant and Globally-Aware Descriptors for Point Cloud Registration

Hao Yu, Ji Hou, Zheng Qin +5

Successful point cloud registration relies on accurate correspondences established upon powerful descriptors. However, existing neural descriptors either leverage a rotation-varian…

cs.CV2018

Markerless Inside-Out Tracking for Interventional Applications

Benjamin Busam, Patrick Ruhkamp, Salvatore Virga +4

Tracking of rotation and translation of medical instruments plays a substantial role in many modern interventions. Traditional external optical tracking systems are often subject t…

cs.CV2020

CPS++: Improving Class-level 6D Pose and Shape Estimation From Monocular Images With Self-Supervised Learning

Fabian Manhardt, Gu Wang, Benjamin Busam +5

Contemporary monocular 6D pose estimation methods can only cope with a handful of object instances. This naturally hampers possible applications as, for instance, robots seamlessly…

cs.CV2026

ODeform: Learning Continuous 4D Motion for Shape Deformation with Neural ODEs

Yordanka Velikova, Mahdi Saleh, Liming Kuang +1

Modeling continuous object deformation is important for many computer vision and robotics tasks, such as manipulation and simulation. Existing approaches rely on learning-based met…

cs.CV2023

On the Importance of Accurate Geometry Data for Dense 3D Vision Tasks

HyunJun Jung, Patrick Ruhkamp, Guangyao Zhai +10

Learning-based methods to solve dense 3D vision problems typically train on 3D sensor data. The respectively used principle of measuring distances provides advantages and drawbacks…

cs.CV2026

GS4City: Hierarchical Semantic Gaussian Splatting via City-Model Priors

Qilin Zhang, Jinyu Zhu, Olaf Wysocki +2

Recent semantic 3D Gaussian Splatting (3DGS) methods primarily rely on 2D foundation models, often yielding ambiguous boundaries and limited support for structured urban semantics.…

cs.CV2020

DynaMiTe: A Dynamic Local Motion Model with Temporal Constraints for Robust Real-Time Feature Matching

Patrick Ruhkamp, Ruiqi Gong, Nassir Navab +1

Feature based visual odometry and SLAM methods require accurate and fast correspondence matching between consecutive image frames for precise camera pose estimation in real-time. C…

cs.CV2022

CroMo: Cross-Modal Learning for Monocular Depth Estimation

Yannick Verdié, Jifei Song, Barnabé Mas +3

Learning-based depth estimation has witnessed recent progress in multiple directions; from self-supervision using monocular video to supervised methods offering highest accuracy. C…

cs.CV2025

Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors

Guangyao Zhai, Yue Zhou, Xinyan Deng +3

Few-shot anomaly detection streamlines and simplifies industrial safety inspection. However, limited samples make accurate differentiation between normal and abnormal features chal…

cs.CV2024

NeRF-Feat: 6D Object Pose Estimation using Feature Rendering

Shishir Reddy Vutukur, Heike Brock, Benjamin Busam +3

Object Pose Estimation is a crucial component in robotic grasping and augmented reality. Learning based approaches typically require training data from a highly accurate CAD model…

cs.CV2025

BioDet: Boosting Industrial Object Detection with Image Preprocessing Strategies

Jiaqi Hu, Hongli Xu, Junwen Huang +3

Accurate 6D pose estimation is essential for robotic manipulation in industrial environments. Existing pipelines typically rely on off-the-shelf object detectors followed by croppi…

cs.CV2023

RIDE: Self-Supervised Learning of Rotation-Equivariant Keypoint Detection and Invariant Description for Endoscopy

Mert Asim Karaoglu, Viktoria Markova, Nassir Navab +2

Unlike in natural images, in endoscopy there is no clear notion of an up-right camera orientation. Endoscopic videos therefore often contain large rotational motions, which require…

cs.CV2020

Project to Adapt: Domain Adaptation for Depth Completion from Noisy and Sparse Sensor Data

Adrian Lopez-Rodriguez, Benjamin Busam, Krystian Mikolajczyk

Depth completion aims to predict a dense depth map from a sparse depth input. The acquisition of dense ground truth annotations for depth completion settings can be difficult and,…

cs.CV2025

EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion

Guangyao Zhai, Evin Pınar Örnek, Dave Zhenyu Chen +5

We present EchoScene, an interactive and controllable generative model that generates 3D indoor scenes on scene graphs. EchoScene leverages a dual-branch diffusion model that dynam…

cs.CV2022

BFS-Net: Weakly Supervised Cell Instance Segmentation from Bright-Field Microscopy Z-Stacks

Shervin Dehghani, Benjamin Busam, Nassir Navab +1

Despite its broad availability, volumetric information acquisition from Bright-Field Microscopy (BFM) is inherently difficult due to the projective nature of the acquisition proces…

cs.CV2022

Disentangling 3D Attributes from a Single 2D Image: Human Pose, Shape and Garment

Xue Hu, Xinghui Li, Benjamin Busam +3

For visual manipulation tasks, we aim to represent image content with semantically meaningful features. However, learning implicit representations from images often lacks interpret…

cs.CV2024

SABER-6D: Shape Representation Based Implicit Object Pose Estimation

Shishir Reddy Vutukur, Mengkejiergeli Ba, Benjamin Busam +2

In this paper, we propose a novel encoder-decoder architecture, named SABER, to learn the 6D pose of the object in the embedding space by learning shape representation at a given p…

cs.CV2023

Polarimetric Information for Multi-Modal 6D Pose Estimation of Photometrically Challenging Objects with Limited Data

Patrick Ruhkamp, Daoyi Gao, HyunJun Jung +2

6D pose estimation pipelines that rely on RGB-only or RGB-D data show limitations for photometrically challenging objects with e.g. textureless surfaces, reflections or transparenc…

cs.CV2022

Bending Graphs: Hierarchical Shape Matching using Gated Optimal Transport

Mahdi Saleh, Shun-Cheng Wu, Luca Cosmo +3

Shape matching has been a long-studied problem for the computer graphics and vision community. The objective is to predict a dense correspondence between meshes that have a certain…

cs.CV2026

MultiCam: On-the-fly Multi-Camera Pose Estimation Using Spatiotemporal Overlaps of Known Objects

Shiyu Li, Hannah Schieber, Kristoffer Waldow +3

Multi-camera dynamic Augmented Reality (AR) applications require a camera pose estimation to leverage individual information from each camera in one common system. This can be achi…

cs.CV2023

3D Adversarial Augmentations for Robust Out-of-Domain Predictions

Alexander Lehner, Stefano Gasperini, Alvaro Marcos-Ramiro +4

Since real-world training datasets cannot properly sample the long tail of the underlying data distribution, corner cases and rare out-of-domain samples can severely hinder the per…

cs.CV2026

3D Segmentation Using Viewpoint-Dependent Spatial Relationships

Ayaka Nanri, Klara Reichard, Mert Kiray +3

Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentation methods do not explicitl…

cs.CV2025

GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation

Weihang Li, Hongli Xu, Junwen Huang +4

A key challenge in model-free category-level pose estimation is the extraction of contextual object features that generalize across varying instances within a specific category. Re…

cs.CV2023

S3M: Scalable Statistical Shape Modeling through Unsupervised Correspondences

Lennart Bastian, Alexander Baumann, Emily Hoppe +5

Statistical shape models (SSMs) are an established way to represent the anatomy of a population with various clinically relevant applications. However, they typically require domai…

cs.CV2024

Zero123-6D: Zero-shot Novel View Synthesis for RGB Category-level 6D Pose Estimation

Francesco Di Felice, Alberto Remus, Stefano Gasperini +5

Estimating the pose of objects through vision is essential to make robotic platforms interact with the environment. Yet, it presents many challenges, often related to the lack of f…

cs.CV2024

SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation

Yamei Chen, Yan Di, Guangyao Zhai +6

Category-level object pose estimation, aiming to predict the 6D pose and 3D size of objects from known categories, typically struggles with large intra-class shape variation. Exist…

cs.CV2023

Segmenting Known Objects and Unseen Unknowns without Prior Knowledge

Stefano Gasperini, Alvaro Marcos-Ramiro, Michael Schmidt +3

Panoptic segmentation methods assign a known class to each pixel given in input. Even for state-of-the-art approaches, this inevitably enforces decisions that systematically lead t…

cs.CV2026

Generative 6D Pose Estimation via Conditional Flow Matching

Amir Hamza, Davide Boscaini, Weihang Li +2

Existing methods for instance-level 6D pose estimation typically rely on neural networks that either directly regress the pose in or estimate it indirectly via loc…

cs.CV2023

TexPose: Neural Texture Learning for Self-Supervised 6D Object Pose Estimation

Hanzhi Chen, Fabian Manhardt, Nassir Navab +1

In this paper, we introduce neural texture learning for 6D object pose estimation from synthetic data and a few unlabelled real images. Our major contribution is a novel learning s…

cs.CV2021

R4Dyn: Exploring Radar for Self-Supervised Monocular Depth Estimation of Dynamic Scenes

Stefano Gasperini, Patrick Koch, Vinzenz Dallabetta +3

While self-supervised monocular depth estimation in driving scenarios has achieved comparable performance to supervised approaches, violations of the static world assumption can st…

cs.CV2021

Attention meets Geometry: Geometry Guided Spatial-Temporal Attention for Consistent Self-Supervised Monocular Depth Estimation

Patrick Ruhkamp, Daoyi Gao, Hanzhi Chen +2

Inferring geometrically consistent dense 3D scenes across a tuple of temporally consecutive images remains challenging for self-supervised monocular depth prediction pipelines. Thi…

cs.CV2026

BenchSeg: A Large-Scale Dataset and Benchmark for Multi-View Food Video Segmentation

Ahmad AlMughrabi, Guillermo Rivo, Carlos Jiménez-Farfán +6

Food image segmentation is a critical task for dietary analysis, enabling accurate estimation of food volume and nutrients. However, current methods suffer from limited multi-view…

cs.CV2025

EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding

Ege Özsoy, Arda Mamur, Felix Tristram +4

Operating rooms (ORs) demand precise coordination among surgeons, nurses, and equipment in a fast-paced, occlusion-heavy environment, necessitating advanced perception models to en…

cs.CV2023

S2P3: Self-Supervised Polarimetric Pose Prediction

Patrick Ruhkamp, Daoyi Gao, Nassir Navab +1

This paper proposes the first self-supervised 6D object pose prediction from multimodal RGB+polarimetric images. The novel training paradigm comprises 1) a physical model to extrac…

cs.CV2025

MS-CLR: Multi-Skeleton Contrastive Learning for Human Action Recognition

Mert Kiray, Alvaro Ritter, Nassir Navab +1

Contrastive learning has gained significant attention in skeleton-based action recognition for its ability to learn robust representations from unlabeled data. However, existing me…

cs.CV2022

Is my Depth Ground-Truth Good Enough? HAMMER -- Highly Accurate Multi-Modal Dataset for DEnse 3D Scene Regression

HyunJun Jung, Patrick Ruhkamp, Guangyao Zhai +9

Depth estimation is a core task in 3D computer vision. Recent methods investigate the task of monocular depth trained with various depth sensor modalities. Every sensor has its adv…

cs.RO2021

ColibriDoc: An Eye-in-Hand Autonomous Trocar Docking System

Shervin Dehghani, Michael Sommersperger, Junjie Yang +6

Retinal surgery is a complex medical procedure that requires exceptional expertise and dexterity. For this purpose, several robotic platforms are currently being developed to enabl…

cs.CV2025

RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation

Junwen Huang, Shishir Reddy Vutukur, Peter KT Yu +3

Typical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the…