activity
20242026
collaborators

5 papers

cs.CV2026

CPPO: Contrastive Perception Policy Optimization for VLM Agents

Ahmad Rezaei, Mohsen Gholami, Saeed Ranjbar Alvar +5

We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core requirement for VLM-based agents…

cs.CV2025

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models

Kevin Cannons, Saeed Ranjbar Alvar, Mohammad Asiful Hossain +6

Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with autonomous driving (AD) being one of…

cs.CV2025

Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes

Mohsen Gholami, Ahmad Rezaei, Zhou Weimin +4

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answeri…

cs.CL2024

Task-Agnostic Language Model Watermarking via High Entropy Passthrough Layers

Vaden Masrani, Mohammad Akbari, David Ming Xuan Yue +2

In the era of costly pre-training of large language models, ensuring the intellectual property rights of model owners, and insuring that said models are responsibly deployed, is be…

cs.CV2024

LaWa: Using Latent Space for In-Generation Image Watermarking

Ahmad Rezaei, Mohammad Akbari, Saeed Ranjbar Alvar +2

With generative models producing high quality images that are indistinguishable from real ones, there is growing concern regarding the malicious usage of AI-generated images. Imper…