activity
20242026
collaborators

6 papers

cs.CV2026

Multimodal Language Models Cannot Spot Spatial Inconsistencies

Om Khangaonkar, Hadi J. Rad, Hamed Pirsiavash

Spatial consistency is a fundamental property of the visual world and a key requirement for models that aim to understand physical reality. Despite recent advances, multimodal larg…

cs.CV2025

Unified Control for Inference-Time Guidance of Denoising Diffusion Models

Maurya Goyal, Anuj Singh, Hadi Jamali-Rad

Aligning diffusion model outputs with downstream objectives is essential for improving task-specific performance. Broadly, inference-time training-free approaches for aligning diff…

cs.LG2025

Graph-Aware Diffusion for Signal Generation

Sergio Rozada, Vimal K. B., Andrea Cavallo +3

We study the problem of generating graph signals from unknown distributions defined over given graphs, relevant to domains such as recommender systems or sensor networks. Our appro…

cs.CV2025

CoDe: Blockwise Control for Denoising Diffusion Models

Anuj Singh, Sayak Mukherjee, Ahmad Beirami +1

Aligning diffusion models to downstream tasks often requires finetuning new models or gradient-based guidance at inference time to enable sampling from the reward-tilted posterior.…

cs.CV2024

MAGMA: Manifold Regularization for MAEs

Alin Dondera, Anuj Singh, Hadi Jamali-Rad

Masked Autoencoders (MAEs) are an important divide in self-supervised learning (SSL) due to their independence from augmentation techniques for generating positive (and/or negative…

cs.CV2024

GeNIe: Generative Hard Negative Images Through Diffusion

Soroush Abbasi Koohpayegani, Anuj Singh, K L Navaneet +2

Data augmentation is crucial in training deep models, preventing them from overfitting to limited data. Recent advances in generative AI, e.g., diffusion models, have enabled more…