9 papers
HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation
An Liu, Bingxi Liu, Hongyu Ding +6
Vision-and-language navigation (VLN) enables robots to follow instructions in previously unseen environments. Recently, a training-free paradigm has emerged: the robot queries a mu…
ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning
Xuchen Liu, Jiawei Huang, Shihao Xia +3
Vision-language navigation (VLN) for UAVs demands grounding free-form instructions into 6-DoF flight under partial observability. While Vision-Language-Action (VLA) models excel at…
MT-PCR: Hybrid Mamba-Transformer Network with Spatial Serialization for Point Cloud Registration
Bingxi Liu, An Liu, Hao Chen +4
Point cloud registration (PCR) is a fundamental task in 3D computer vision and robotics. Most learning-based PCR methods rely on Transformer architectures, which suffer from quadra…
Hierarchical Visual Relocalization with Nearest View Synthesis from Feature Gaussian Splatting
Huaqi Tao, Bingxi Liu, Guangcheng Chen +3
Visual relocalization is a fundamental task in the field of 3D computer vision, estimating a camera's pose when it revisits a previously known scene. While point-based hierarchical…
Learnable Query Aggregation with KV Routing for Cross-view Geo-localisation
Hualin Ye, Bingxi Liu, Jixiang Du +3
Cross-view geo-localisation (CVGL) aims to estimate the geographic location of a query image by matching it with images from a large-scale database. However, the significant view-p…
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
Huaqi Tao, Bingxi Liu, Calvin Chen +4
Visual Place Recognition (VPR) is a crucial capability for long-term autonomous robots, enabling them to identify previously visited locations using visual information. However, ex…