2 citations · 2 across the 4 of their papers we have counts for
4 papers
Smooth Operator: Smooth Verifiable Reward Activates Spatial Reasoning Ability of Vision-Language Model
Siwen Jiao, Tianxiong Lv, Kangan Qian +7
Vision-Language Models (VLMs) face a critical bottleneck in achieving precise numerical prediction for 3D scene understanding. Traditional reinforcement learning (RL) approaches, p…
EvaDrive: Evolutionary Adversarial Policy Optimization for End-to-End Autonomous Driving
Siwen Jiao, Kangan Qian, Hao Ye +12
Autonomous driving faces significant challenges in achieving human-like iterative decision-making, which continuously generates, evaluates, and refines trajectory proposals. Curren…
A Survey on Vision-Language-Action Models for Autonomous Driving
Sicong Jiang, Zilin Huang, Kangan Qian +17
The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language unde…
LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement
Siwen Jiao, Yangyi Fang, Baoyun Peng +2
Recent advancements in Visual Language Models (VLMs) have made them crucial for visual question answering (VQA) in autonomous driving, enabling natural human-vehicle interactions.…