9 papers
GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors
Jiazheng Liu, Hang Li, Jiawei Zhang +5
Generative models like Diffusion Models and Flow Matching have demonstrated remarkable capabilities in synthesizing high-fidelity driving videos, but are severely constrained by hi…
Learning Gaussian Structure: Intervention-Guided Density Control for Feed-Forward Driving Reconstruction
Hang Li, Jiahe Li, Meiying Gu +3
Feed-forward Gaussian reconstruction has recently emerged as an efficient approach for driving scene reconstruction. However, prevailing LiDAR-based methods preserve the initial co…
Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction
Jiahe Li, Jiawei Zhang, Xiao Bai +4
Surface reconstruction with differentiable rendering has achieved impressive performance in recent years, yet the pervasive photometric ambiguities have strictly bottlenecked exist…
Demystifying Action Space Design for Robotic Manipulation Policies
Yuchun Feng, Jinliang Zheng, Zhihao Wang +5
The specification of the action space plays a pivotal role in imitation-based robotic manipulation policy learning, fundamentally shaping the optimization landscape of policy learn…
SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate Embeddings
Yuchen Wu, Jiahe Li, Xiaohan Yu +3
Monocular visual SLAM enables 3D reconstruction from internet video and autonomous navigation on resource-constrained platforms, yet suffers from scale drift, i.e., the gradual div…
FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM
Yuchen Wu, Jiahe Li, Fabio Tosi +3
We present FoundationSLAM, a learning-based monocular dense SLAM system that addresses the absence of geometric consistency in previous flow-based approaches for accurate and robus…