6 papers
Multimodal Language Models Cannot Spot Spatial Inconsistencies
Om Khangaonkar, Hadi J. Rad, Hamed Pirsiavash
Spatial consistency is a fundamental property of the visual world and a key requirement for models that aim to understand physical reality. Despite recent advances, multimodal larg…
Unified Control for Inference-Time Guidance of Denoising Diffusion Models
Maurya Goyal, Anuj Singh, Hadi Jamali-Rad
Aligning diffusion model outputs with downstream objectives is essential for improving task-specific performance. Broadly, inference-time training-free approaches for aligning diff…
Graph-Aware Diffusion for Signal Generation
Sergio Rozada, Vimal K. B., Andrea Cavallo +3
We study the problem of generating graph signals from unknown distributions defined over given graphs, relevant to domains such as recommender systems or sensor networks. Our appro…
CoDe: Blockwise Control for Denoising Diffusion Models
Anuj Singh, Sayak Mukherjee, Ahmad Beirami +1
Aligning diffusion models to downstream tasks often requires finetuning new models or gradient-based guidance at inference time to enable sampling from the reward-tilted posterior.…
MAGMA: Manifold Regularization for MAEs
Alin Dondera, Anuj Singh, Hadi Jamali-Rad
Masked Autoencoders (MAEs) are an important divide in self-supervised learning (SSL) due to their independence from augmentation techniques for generating positive (and/or negative…
GeNIe: Generative Hard Negative Images Through Diffusion
Soroush Abbasi Koohpayegani, Anuj Singh, K L Navaneet +2
Data augmentation is crucial in training deep models, preventing them from overfitting to limited data. Recent advances in generative AI, e.g., diffusion models, have enabled more…