34 papers
SeqLoc: Beyond the Single Frame for Cross-View Geo-Localization in Feature-Sparse Scenes
Junwei Zheng, Yun Huang, Ruize Dai +8
Cross-View Geo-Localization (CVGL) with OpenStreetMap (OSM) performs well in structure-rich urban environments but collapses in feature-sparse scenes such as rural roads. To study…
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
Ruiping Liu, Junwei Zheng, Yufan Chen +7
Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial-temporal reasoning. Inspired by human recall from both field and obs…
What if? Emulative Simulation with World Models for Situated Reasoning
Ruiping Liu, Yufan Chen, Yuheng Zhang +8
Situated reasoning often relies on active exploration, yet in many real-world scenarios such exploration is infeasible due to physical constraints of robots or safety concerns of v…
InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing
Yebin Yang, Di Wen, Lei Qi +10
Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity…
RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization
Junwei Zheng, Ruize Dai, Ruiping Liu +7
Metric Cross-View Geo-Localization (MCVGL) aims to estimate the 3-DoF camera pose (position and heading) by matching ground and satellite images. In this work, instead of pinhole a…
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
Kunyu Peng, Zhikun Zhou, Kailun Yang +9
Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoin…