6 papers
Enhancing Policy Learning with World-Action Model
Yuci Han, Alper Yilmaz
This paper presents the World-Action Model (WAM), an action-regularized world model that jointly reasons over future visual observations and the actions that drive state transition…
Zero-shot Vision-Language Reranking for Cross-View Geolocalization
Yunus Talha Erzurumlu, John E. Anderson, William J. Shuart +2
Cross-view geolocalization (CVGL) systems, while effective at retrieving a list of relevant candidates (high Recall@k), often fail to identify the single best match (low Top-1 accu…
Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming
Yunus Talha Erzurumlu, Jiyong Kwag, Alper Yilmaz
Cross-view geo-localization (CVGL) estimates a camera's location by matching a street-view image to geo-referenced overhead imagery, enabling GPS-denied localization and navigation…
BetterScene: 3D Scene Synthesis with Representation-Aligned Generative Model
Yuci Han, Charles Toth, John E. Anderson +2
We present BetterScene, an approach to enhance novel view synthesis (NVS) quality for diverse real-world scenes using extremely sparse, unconstrained photos. BetterScene leverages…
Lightweight Road Environment Segmentation using Vector Quantization
Jiyong Kwag, Alper Yilmaz, Charles Toth
Road environment segmentation plays a significant role in autonomous driving. Numerous works based on Fully Convolutional Networks (FCNs) and Transformer architectures have been pr…
UAS Visual Navigation in Large and Unseen Environments via a Meta Agent
Yuci Han, Charles Toth, Alper Yilmaz
The aim of this work is to develop an approach that enables Unmanned Aerial System (UAS) to efficiently learn to navigate in large-scale urban environments and transfer their acqui…