6 papers
SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic Reasoning
Furong Jia, Ling Dai, Wenjin Deng +4
Large Vision-Language Models (LVLMs) have demonstrated strong reasoning capabilities in geo-localization, yet they often struggle in real-world scenarios where visual cues are spar…
Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training
Wenshuo Wang, Fan Zhang
Zero-Shot Super-Resolution Spatiotemporal Forecasting requires a deep learning model to be trained on low-resolution data and deployed for inference on high-resolution. Existing st…
GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning
Rui Tang, Guankun Wang, Long Bai +6
Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains hig…
CUS-GS: A Compact Unified Structured Gaussian Splatting Framework for Multimodal Scene Representation
Yuhang Ming, Chenxin Fang, Xingyuan Yu +4
Recent advances in Gaussian Splatting based 3D scene representation have shown two major trends: semantics-oriented approaches that focus on high-level understanding but lack expli…
Towards Interpretable Geo-localization: a Concept-Aware Global Image-GPS Alignment Framework
Furong Jia, Lanxin Liu, Ce Hou +3
Worldwide geo-localization involves determining the exact geographic location of images captured globally, typically guided by geographic cues such as climate, landmarks, and archi…
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
Ruizhe Ou, Yuan Hu, Fan Zhang +2
Multi-modal large language models (MLLMs) have achieved remarkable success in image- and region-level remote sensing (RS) image understanding tasks, such as image captioning, visua…