6 papers
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
Karim Knaebel, Kadir Yilmaz, Daan de Geus +4
Vision foundation models (VFMs) trained on large-scale image datasets provide high-quality features that have significantly advanced 2D visual recognition. However, their potential…
How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction?
Stephanie Käs, Anton Burenko, Louis Markert +4
Gestures enable non-verbal human-robot communication, especially in noisy environments like agile production. Traditional deep learning-based gesture recognition relies on task-spe…
Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images
Stephanie Käs, Sven Peter, Henrik Thillmann +5
Fisheye cameras offer robots the ability to capture human movements across a wider field of view (FOV) than standard pinhole cameras, making them particularly useful for applicatio…
OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
Jens Piekenbrinck, Christian Schmidt, Alexander Hermans +3
3D Gaussian Splatting (3DGS) has emerged as a powerful representation for neural scene reconstruction, offering high-quality novel view synthesis while maintaining computational ef…
UPTor: Unified 3D Human Pose Dynamics and Trajectory Prediction for Human-Robot Interaction
Nisarga Nilavadi, Andrey Rudenko, Timm Linder
We introduce a unified approach to forecast the dynamics of human keypoints along with the motion trajectory based on a short sequence of input poses. While many studies address ei…
Acquisition of high-quality images for camera calibration in robotics applications via speech prompts
Timm Linder, Kadir Yilmaz, David B. Adrian +1
Accurate intrinsic and extrinsic camera calibration can be an important prerequisite for robotic applications that rely on vision as input. While there is ongoing research on enabl…