4 papers
Embody4D: A Generalist Data Engine for Embodied 4D World Modeling
Peiyan Tu, Hanxin Zhu, Jingwen Sun +6
Embodied agents require robust and comprehensive 3D spatiotemporal representations to support spatial reasoning, manipulation understanding, and downstream decision making. However…
Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment
Cong Wang, Hanxin Zhu, Jiayi Luo +6
Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coherent and visually convincin…
ViewFormer: Exploring Spatiotemporal Modeling for Multi-View 3D Occupancy Perception via View-Guided Transformers
Jinke Li, Xiao He, Chonghua Zhou +3
3D occupancy, an advanced perception technology for driving scenarios, represents the entire scene without distinguishing between foreground and background by quantifying the physi…
The RoboDrive Challenge: Drive Anytime Anywhere in Any Condition
Lingdong Kong, Shaoyuan Xie, Hanjiang Hu +88
In the realm of autonomous driving, robust perception under out-of-distribution conditions is paramount for the safe deployment of vehicles. Challenges such as adverse weather, sen…