6 papers
DAP: A Discrete-token Autoregressive Planner for Autonomous Driving
Bowen Ye, Bin Zhang, Hang Zhao
Gaining sustainable performance improvement with scaling data and model budget remains a pivotal yet unresolved challenge in autonomous driving. While autoregressive models exhibit…
ActionCodec: What Makes for Good Action Tokenizers
Zibin Dong, Yicheng Liu, Shiduo Zhang +8
Vision-Language-Action (VLA) models leveraging the native autoregressive paradigm of Vision-Language Models (VLMs) have demonstrated superior instruction-following and training eff…
Embodied Intelligence for Flexible Manufacturing: A Survey
Kai Xu, Hang Zhao, Ruizhen Hu +4
Driven by breakthroughs in next-generation artificial intelligence, embodied intelligence is rapidly advancing into industrial manufacturing. In flexible manufacturing, industrial…
SAMP: Spatial Anchor-based Motion Policy for Collision-Aware Robotic Manipulators
Kai Chen, Zhihai Bi, Guoyang Zhao +4
Neural-based motion planning methods have achieved remarkable progress for robotic manipulators, yet a fundamental challenge lies in simultaneously accounting for both the robot's…
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
Ruixun Liu, Lingyu Kong, Derun Li +1
Multimodal large language models (MLLMs) have shown strong vision-language reasoning abilities but still lack robust 3D spatial understanding, which is critical for autonomous driv…
BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving
Zeming Chen, Hang Zhao
Multi-view image generation in autonomous driving demands consistent 3D scene understanding across camera views. Most existing methods treat this problem as a 2D image set generati…