activity
20242026
collaborators

5 papers

cs.CV2026

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models

Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain +6

Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of…

cs.CV2026

CPPO: Contrastive Perception Policy Optimization for VLM Agents

Ahmad Rezaei, Mohsen Gholami, Saeed Ranjbar Alvar +5

We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core requirement for VLM-based agents…

cs.CV2025

Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes

Mohsen Gholami, Ahmad Rezaei, Zhou Weimin +4

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answeri…

cs.CV2025

LaWa: Using Latent Space for In-Generation Image Watermarking

Ahmad Rezaei, Mohammad Akbari, Saeed Ranjbar Alvar +2

With generative models producing high quality images that are indistinguishable from real ones, there is growing concern regarding the malicious usage of AI-generated images. Imper…

cs.CL2024

Task-Agnostic Language Model Watermarking via High Entropy Passthrough Layers

Vaden Masrani, Mohammad Akbari, David Ming Xuan Yue +2

In the era of costly pre-training of large language models, ensuring the intellectual property rights of model owners, and insuring that said models are responsibly deployed, is be…