12 papers
InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving
Xiaoyu Ye, Leheng Li, Xinyu Ji +8
Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introdu…
Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
Jing He, Haodong Li, Mingzhi Sheng +1
Recovering pixel-wise geometric properties from a single image is fundamentally ill-posed due to appearance ambiguity and non-injective mappings between 2D observations and 3D stru…
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
Haodong Yan, Zhide Zhong, Jiaguan Zhu +10
Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs,…
BuildAnyPoint: 3D Building Structured Abstraction from Diverse Point Clouds
Tongyan Hua, Haoran Gong, Yuan Liu +3
We introduce BuildAnyPoint, a novel generative framework for structured 3D building reconstruction from point clouds with diverse distributions, such as those captured by airborne…
FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control
Mingzhi Sheng, Zekai Gu, Peng Li +4
Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task-specific signals, we argue that a fundamental…
SPOT-Occ: Sparse Prototype-guided Transformer for Camera-based 3D Occupancy Prediction
Suzeyu Chen, Leheng Li, Ying-Cong Chen
Achieving highly accurate and real-time 3D occupancy prediction from cameras is a critical requirement for the safe and practical deployment of autonomous vehicles. While this shif…