6 papers
GeoWorldAD: Geometry World Action Model for Autonomous Driving
Songyan Zhang, Jinyuan Tian, Hanbing Li +9
Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual ob…
Structure-Detail Decoupled Autoregressive Generation for Fast and High-Fidelity Virtual Try-On
Lu Yang, Xiaonan Hu, Yanan Li +3
The paper introduces STAR-VTON, a two‑stage autoregressive framework for virtual try‑on that generates garment structure in a latent space and then refines fine‑grained details in…
Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation
Hong Chen, Daqi Liu, Zehan Zhang +10
Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, decoupling goal intent from trajectory synthesis. This approach suffers…
AutoMine Solution for AV2 2026 Scenario Mining Challenge
Songliang Cao, Jiele Zhao, Yuru Wang +10
With the development of autonomous driving systems, mining high-value, safety-critical, and planning-relevant scenarios from large-scale driving logs has become essential for data-…
MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving
Haiguang Wang, Daqi Liu, Hongwei Xie +5
In recent years, data-driven techniques have greatly advanced autonomous driving systems, but the need for rare and diverse training data remains a challenge, requiring significant…
Learning A Zero-shot Occupancy Network from Vision Foundation Models via Self-supervised Adaptation
Sihao Lin, Daqi Liu, Ruochong Fu +6
Estimating the 3D world from 2D monocular images is a fundamental yet challenging task due to the labour-intensive nature of 3D annotations. To simplify label acquisition, this wor…