3 papers
cs.CV2026
Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting
Panav Shah, Geet Sethi, Ashutosh Gandhe
Visual grounding aims to locate image regions that correspond to natural language descriptions and is a key component of interpretable vision systems. In remote sensing imagery, gr…
cs.CV2026
DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery
Geet Sethi, Panav Shah, Ashutosh Gandhe +1
Diffusion models have emerged as powerful tools for a wide range of vision tasks, including text-guided image generation and editing. In this work, we explore their potential for o…
cs.CV2025
Movie Gen: A Cast of Media Foundation Models
Adam Polyak, Amit Zohar, Andrew Brown +85
We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabili…