11 papers
Driving Like Yourself: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving
Xiaoru Dong, Ruiqin Li, Xiao Han +7
Human driving behavior is inherently diverse, yet most end-to-end autonomous driving (E2E-AD) systems learn a single average driving style, neglecting individual differences. Achie…
From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges
Yiming Zhong, Yaoyu He, Zemin Yang +5
Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal sca…
Implicit Drifting Policy: One-Step Action Generation via Conditional Expert Geometry
Zemin Yang, Yaoyu He, Yiming Zhong +5
Generative action policies based on diffusion or flow matching excel in behavior cloning, yet their iterative sampling is prohibitive for high-frequency robot control. While recent…
Controllable Video Object Insertion via Multi-View Priors
Qi Xia, Xia Qi, Peishan Cong +4
Video object insertion places a user-specified object in an existing dynamic scene. Existing methods typically condition generation on text or a single reference image. Consequentl…
HUMOF: Human Motion Forecasting in Interactive Social Scenes
Caiyi Sun, Yujing Sun, Xiao Han +5
Complex scenes present significant challenges for predicting human behaviour due to the abundance of interaction information, such as human-human and humanenvironment interactions.…
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
Yu Li, Yuenan Hou, Yingmei Wei +4
Multi-modal 3D understanding is a fundamental task in computer vision. Previous multi-modal fusion methods typically employ a single, dense fusion network, struggling to handle the…