23 citations · 53 across the 64 of their papers we have counts for
11 papers · 1 filter
What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency
Luoyang Sun, Guoyang Xia, Fengfa Li +9
Vision-Language-Action (VLA) models combine a pretrained vision encoder, a language backbone, and an action head, but their relative contribution has not been established under con…
BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
Bing Zhan, Shuyao Shang, Shuo Lu +6
Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of…
ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving
Huimin Wang, Yue Wang, Bihao Cui +7
We introduce ReflectDrive-2, a masked discrete diffusion planner with separate action expert for autonomous driving that represents plans as discrete trajectory tokens and generate…
WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving
Pengxuan Yang, Ben Lu, Zhongpu Xia +7
Latent World Models enhance scene representation through temporal self-supervised learning, presenting a perception annotation-free paradigm for end-to-end autonomous driving. Howe…
Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
Pengxiang Li, Yinan Zheng, Yue Wang +6
End-to-End (E2E) solutions have emerged as a mainstream approach for autonomous driving systems, with Vision-Language-Action (VLA) models representing a new paradigm that leverages…
The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
Titong Jiang, Xuefeng Jiang, Yuan Ma +7
We present LightVLA, a simple yet effective differentiable token pruning framework for vision-language-action (VLA) models. While VLA models have shown impressive capability in exe…