12 papers
TerraDiT-: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive
Brian Wei, Srikumar Sastry, Daniel Cher +2
Generative models have achieved remarkable progress, yet applying them to satellite imagery remains challenging. Unlike natural imagery, satellite scenes are structured by spatiall…
TerraDiT: Point-Conditioned Diffusion Transformer for Satellite Image Synthesis
Srikumar Sastry, Dan Cher, Brian Wei +4
We introduce TerraDiT, a diffusion transformer designed for text-to-satellite image generation with point-based control. Existing controlled satellite image generative models often…
Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification
Zhiyuan Tao, Srikumar Sastry, Matthew J Thompson +9
Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically structured label spaces, existing…
ENPIRE: Agentic Robot Policy Self-Improvement in the Real World
Wenli Xiao, Jia Xie, Tonghe Zhang +14
Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of gener…
DiffVAS: Diffusion-Guided Visual Active Search in Partially Observable Environments
Anindya Sarkar, Srikumar Sastry, Aleksis Pirinen +2
Visual active search (VAS) has been introduced as a modeling framework that leverages visual cues to direct aerial (e.g., UAV-based) exploration and pinpoint areas of interest with…
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping
Subash Khanal, Srikumar Sastry, Aayush Dhakal +3
We present Sat2Sound, a unified multimodal framework for geospatial soundscape understanding, designed to predict and map the distribution of sounds across the Earth's surface. Exi…