5 papers
Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation
Hong Chen, Daqi Liu, Zehan Zhang +10
Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, decoupling goal intent from trajectory synthesis. This approach suffers…
Arbitrary Generative Video Interpolation
Guozhen Zhang, Haiguang Wang, Chunyu Wang +3
Video frame interpolation (VFI), which generates intermediate frames from given start and end frames, has become a fundamental function in video generation applications. However, e…
MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving
Haiguang Wang, Daqi Liu, Hongwei Xie +5
In recent years, data-driven techniques have greatly advanced autonomous driving systems, but the need for rare and diverse training data remains a challenge, requiring significant…
Fully Sparse 3D Occupancy Prediction
Haisong Liu, Yang Chen, Haiguang Wang +6
Occupancy prediction plays a pivotal role in autonomous driving. Previous methods typically construct dense 3D volumes, neglecting the inherent sparsity of the scene and suffering…
LAIP: Learning Local Alignment from Image-Phrase Modeling for Text-based Person Search
Haiguang Wang, Yu Wu, Mengxia Wu +2
Text-based person search aims at retrieving images of a particular person based on a given textual description. A common solution for this task is to directly match the entire imag…