activity
20172026
most citedDriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

23 citations · 53 across the 69 of their papers we have counts for

collaborators
Showing 2026 · cs.ROShow all

5 papers · 2 filters

cs.RO2026

MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence

Haoran Wen, Wenfu Wang, Kunsong Shi +15

General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-ac…

cs.RO2026

ME-Brain-1.0: Memory, Cognition and Action for Evolving Embodied Intelligence

Wei He, Hengtao Li, Zhongrui Yu +20

Current embodied systems largely rely on pretrained capabilities that remain fixed after deployment, limiting their ability to learn from physical interaction. We introduce MachEmb…

cs.RO2026

What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency

Luoyang Sun, Guoyang Xia, Fengfa Li +9

Vision-Language-Action (VLA) models combine a pretrained vision encoder, a language backbone, and an action head, but their relative contribution has not been established under con…

cs.RO2026

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

Bing Zhan, Shuyao Shang, Shuo Lu +6

Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of…

cs.RO2026

ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving

Huimin Wang, Yue Wang, Bihao Cui +7

We introduce ReflectDrive-2, a masked discrete diffusion planner with separate action expert for autonomous driving that represents plans as discrete trajectory tokens and generate…