8 papers
Who Responds When the Driver Is Gone? A Framework for Holistic Passenger Intent Understanding
Xuewen Luo, Ding Fan, Ruiqi Chen +5
As autonomous vehicles advance toward driverless mobility, understanding and responding to passenger needs and intentions becomes increasingly important in the absence of a human d…
CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models
Fengze Yang, Bo Yu, Xuewen Luo +2
Vision-Language Models (VLMs) face severe memory and latency bottlenecks due to high-resolution visual tokens. While current token reduction methods theoretically save FLOPs, post-…
KnowDiffuser: A Knowledge-Guided Diffusion Planner with LLM Reasoning
Fan Ding, Xuewen Luo, Fengze Yang +4
Recent advancements in Language Models (LMs) have demonstrated strong semantic reasoning capabilities, enabling their application in high-level decision-making for autonomous drivi…
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
Bo Yu, Fengze Yang, Yiming Liu +6
The emergence of Vision-Language Models (VLMs) has introduced new paradigms for global image geo-localization through retrieval-augmented generation (RAG) and reasoning-driven infe…
V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving
Xuewen Luo, Fengze Yang, Fan Ding +5
Autonomous driving (AD) has achieved significant progress, yet single-vehicle perception remains constrained by sensing range and occlusions. Vehicle-to-Everything (V2X) communicat…
Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
Fengze Yang, Bo Yu, Yang Zhou +3
Autonomous driving (AD) systems relying solely on onboard sensors may fail to detect distant or obstacle hazards, potentially causing preventable collisions; however, existing tran…