12 papers
P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture
Felix Tristram, Stefano Gasperini, Benjamin Killeen +4
The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assistance systems for complex, mu…
Prompting Diffusion Models for Zero-Shot Instance Segmentation
Irem Zeynep Alagöz, Nils Morbitzer, Andrea Ramazzina +3
Several disruptive research directions have recently emerged in computer vision, including foundation models achieving previously unseen zero-shot performance in scene understandin…
Future Dynamic 3D Reconstruction: Toward 3D World Modeling with Disentangled Ego-Motion
Nils Morbitzer, Jonathan Evers, Artem Savkin +4
Forecasting the evolution of dynamic environments is crucial for autonomous agents. While generative world models have achieved high photorealism in 2D video synthesis by mixing eg…
SA4Depth: Consistent Pose-Depth Scale Alignment for Self-Supervised Monocular Depth Estimation
Changxuan Li, Nadine Berner, Nassir Navab +2
Self-supervised depth estimation from monocular sequences relies on the joint learning of a depth and a pose network. Despite abundant research done to improve the depth network, e…
OpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook Attention
Kunyi Li, Michael Niemeyer, Sen Wang +3
Understanding open-vocabulary 3D scenes with Gaussian-based representations remains challenging due to fragmented and spatially inconsistent semantic predictions across multi-view…
SING3R-SLAM: Submap-based Indoor Monocular Gaussian SLAM with 3D Reconstruction Priors
Kunyi Li, Michael Niemeyer, Sen Wang +3
Recent advances in dense 3D reconstruction have demonstrated strong capability in accurately capturing local geometry. However, extending these methods to incremental global recons…