3 papers
cs.CV2026
From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models
Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain +6
Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of…
cs.CV2026
CPPO: Contrastive Perception Policy Optimization for VLM Agents
Ahmad Rezaei, Mohsen Gholami, Saeed Ranjbar Alvar +5
We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core requirement for VLM-based agents…
cs.CV2025
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
Mohsen Gholami, Ahmad Rezaei, Zhou Weimin +4
Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answeri…