5 papers
Point Prompting: Counterfactual Tracking with Video Diffusion Models
Ayush Shrivastava, Sanyam Mehta, Daniel Geng +1
Trackers and video generators solve closely related problems: the former analyze motion, while the latter synthesize it. We show that this connection enables pretrained video diffu…
ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery
Ayush Shrivastava, Kirtan Gangani, Laksh Jain +2
Vision language models (VLMs) achieve strong performance on RGB imagery, but they do not generalize to thermal images. Thermal sensing plays a critical role in settings where visib…
Fine-grained Defocus Blur Control for Generative Image Models
Ayush Shrivastava, Connelly Barnes, Xuaner Zhang +4
Current text-to-image diffusion models excel at generating diverse, high-quality images, yet they struggle to incorporate fine-grained camera metadata such as precise aperture sett…
Self-Supervised Spatial Correspondence Across Modalities
Ayush Shrivastava, Andrew Owens
We present a method for finding cross-modal space-time correspondences. Given two images from different visual modalities, such as an RGB image and a depth map, our model identifie…
Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation
Rohith Peddi, Saurabh, Ayush Abhay Shrivastava +2
Spatio-Temporal Scene Graphs (STSGs) provide a concise and expressive representation of dynamic scenes by modeling objects and their evolving relationships over time. However, real…