3 papers
cs.CV2026
Point Prompting: Counterfactual Tracking with Video Diffusion Models
Ayush Shrivastava, Sanyam Mehta, Daniel Geng +1
Trackers and video generators solve closely related problems: the former analyze motion, while the latter synthesize it. We show that this connection enables pretrained video diffu…
cs.CV2025
Fine-grained Defocus Blur Control for Generative Image Models
Ayush Shrivastava, Connelly Barnes, Xuaner Zhang +4
Current text-to-image diffusion models excel at generating diverse, high-quality images, yet they struggle to incorporate fine-grained camera metadata such as precise aperture sett…
cs.CV2025
Self-Supervised Spatial Correspondence Across Modalities
Ayush Shrivastava, Andrew Owens
We present a method for finding cross-modal space-time correspondences. Given two images from different visual modalities, such as an RGB image and a depth map, our model identifie…