8 papers
Latent World Models with Monotone Planning Costs for Image-Goal Navigation
Amirhosein Chahe, Siwei Cai, Lifeng Zhou
Image-goal navigation with latent world models requires not only accurate future prediction, but also a planning cost that reliably ranks candidate action sequences. We define the…
What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models
Amirhosein Chahe, Tyler Naes, Jovin D'sa +4
Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view. Current approaches often treat all occlusions with uniform con…
Fuzzy Encoding-Decoding to Improve Spiking Q-Learning Performance in Autonomous Driving
Aref Ghoreishee, Abhishek Mishra, Lifeng Zhou +3
This paper develops an end-to-end fuzzy encoder-decoder architecture for enhancing vision-based multi-modal deep spiking Q-networks in autonomous driving. The method addresses two…
IDSelect: A RL-Based Cost-Aware Selection Agent for Video-based Multi-Modal Person Recognition
Yuyang Ji, Yixuan Shen, Kien Nguyen +2
Video-based person recognition achieves robust identification by integrating face, body, and gait. However, current systems waste computational resources by processing all modaliti…
New Spiking Architecture for Multi-Modal Decision-Making in Autonomous Vehicles
Aref Ghoreishee, Abhishek Mishra, Lifeng Zhou +2
This work proposes an end-to-end multi-modal reinforcement learning framework for high-level decision-making in autonomous vehicles. The framework integrates heterogeneous sensory…
ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
Amirhosein Chahe, Lifeng Zhou
Vision-language models (VLMs) show promise for autonomous driving but often lack transparent reasoning capabilities that are critical for safety. We investigate whether explicitly…