2 papers
cs.CV2026
STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision
Pou-Chun Kung, Aryaman Rao, Utkrisht Sahai +4
Vision-language models (VLMs) are emerging as a key component of embodied intelligence, with growing applications in auto-labeling and end-to-end autonomous driving. However, exist…
cs.CL2026
EVLA: An Electro-Aware Multimodal Assistant for Physically-Grounded Driving Reasoning and Control
Yuxin Liu, Zihan Chen, Haoyu Wang +3
Modern vision-language models (VLMs) for driving assistants typically treat vehicle dynamics as a black box, resulting in decisions that lack awareness of the vehicle's real-time e…