7 papers
AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
Guiyu Zhao, Longteng Guo, Yanghong Mei +7
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon task…
NavWM: A Unified Navigation World Model for Foresight-Driven Planning
Yanghong Mei, Longteng Guo, Ming-Ming Yu +3
Conventional visual navigation policies often struggle with myopic decision-making and mode collapse in complex environments. While world models offer a promising alternative, exis…
VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models
Guiyu Zhao, Longteng Guo, Junyou Zhu +6
Vision-language-action (VLA) models have shown strong promise for robotic manipulation, but their reliability at test time remains limited by one-shot action prediction, where even…
Cross-Layer Feature Pyramid Transformer for Small Object Detection in Aerial Images
Zewen Du, Zhenjiang Hu, Guiyu Zhao +2
Object detection in aerial images has always been a challenging task due to the generally small size of the objects. Most current detectors prioritize the development of new detect…
Progressive Correspondence Regenerator for Robust 3D Registration
Guiyu Zhao, Sheng Ao, Ye Zhang +2
Obtaining enough high-quality correspondences is crucial for robust registration. Existing correspondence refinement methods mostly follow the paradigm of outlier removal, which ei…
Cross-PCR: A Robust Cross-Source Point Cloud Registration Framework
Guiyu Zhao, Zhentao Guo, Zewen Du +1
Due to the density inconsistency and distribution difference between cross-source point clouds, previous methods fail in cross-source point cloud registration. We propose a density…