activity
20242026
collaborators

8 papers

cs.CV2026

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

Jing Wu, Jianhua Wu, Jiayi Guan +5

Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or ext…

cs.CV2026

ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving

Yujin Wang, Yutong Zheng, Wenxian Fan +7

In this paper, we introduce ScenePilot-4K, a large-scale first-person dataset for safety-aware vision-language learning and evaluation in autonomous driving. Built from public onli…

cs.CV2025

RAC3: Retrieval-Augmented Corner Case Comprehension for Autonomous Driving with Vision-Language Models

Yujin Wang, Quanfeng Liu, Jiaqi Fan +5

Understanding and addressing corner cases is essential for ensuring the safety and reliability of autonomous driving systems. Vision-language models (VLMs) play a crucial role in e…

cs.RO2025

Relative Position Matters: Trajectory Prediction and Planning with Polar Representation

Bozhou Zhang, Nan Song, Bingzhao Gao +1

Trajectory prediction and planning in autonomous driving are highly challenging due to the complexity of predicting surrounding agents' movements and planning the ego agent's actio…

cs.RO2025

BIDA: A Bi-level Interaction Decision-making Algorithm for Autonomous Vehicles in Dynamic Traffic Scenarios

Liyang Yu, Tianyi Wang, Junfeng Jiao +3

In complex real-world traffic environments, autonomous vehicles (AVs) need to interact with other traffic participants while making real-time and safety-critical decisions accordin…

cs.CV2025

RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving

Yujin Wang, Quanfeng Liu, Zhengxin Jiang +5

Accurately understanding and deciding high-level meta-actions is essential for ensuring reliable and safe autonomous driving systems. While vision-language models (VLMs) have shown…