6 papers
BA-T: An Iterative Transformer for Two-View Bundle Adjustment
Ganlin Zhang, Weirong Chen, Daniel Cremers +1
Feed-forward models for 3D reconstruction have achieved strong performance using deep cross-view attention to exchange information across images. However, these approaches often de…
EgoSpot:Egocentric Multimodal Control for Hands-Free Mobile Manipulation
Ganlin Zhang, Deheng Zhang, Longteng Duan +5
We propose a novel hands-free control framework for the Boston Dynamics Spot robot using the Microsoft HoloLens 2 mixed-reality headset. Enabling accessible robot control is critic…
NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
Weirong Chen, Chuanxia Zheng, Ganlin Zhang +2
We present NOVA3R, an effective approach for non-pixel-aligned 3D reconstruction from a set of unposed images in a feed-forward manner. Unlike pixel-aligned methods that tie geomet…
Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
Shenhan Qian, Ganlin Zhang, Shangzhe Wu +1
Reconstructing and tracking dynamic 3D scenes is a fundamental challenge in computer vision. Existing methods typically decouple geometry from motion: static multi-view reconstruct…
ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association
Ganlin Zhang, Shenhan Qian, Xi Wang +1
We present ViSTA-SLAM as a real-time monocular visual SLAM system that operates without requiring camera intrinsics, making it broadly applicable across diverse camera setups. At i…
Back on Track: Bundle Adjustment for Dynamic Scene Reconstruction
Weirong Chen, Ganlin Zhang, Felix Wimbauer +4
Traditional SLAM systems, which rely on bundle adjustment, struggle with highly dynamic scenes commonly found in casual videos. Such videos entangle the motion of dynamic elements,…