1 citations · 1 across the 8 of their papers we have counts for
11 papers
iGaussian: Real-Time Camera Pose Estimation via Feed-Forward 3D Gaussian Splatting Inversion
Hao Wang, Linqing Zhao, Xiuwei Xu +2
Recent trends in SLAM and visual navigation have embraced 3D Gaussians as the preferred scene representation, highlighting the importance of estimating camera poses from a single i…
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
Hang Yin, Haoyu Wei, Xiuwei Xu +3
In this paper, we propose a training-free framework for vision-and-language navigation (VLN). Existing zero-shot VLN methods are mainly designed for discrete environments or involv…
MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation
Zhenyu Wu, Angyuan Ma, Xiuwei Xu +5
Mobile manipulation stands as a core challenge in robotics, enabling robots to assist humans across varied tasks and dynamic daily environments. Conventional mobile manipulation ap…
IGL-Nav: Incremental 3D Gaussian Localization for Image-goal Navigation
Wenxuan Guo, Xiuwei Xu, Hang Yin +4
Visual navigation with an image as goal is a fundamental and challenging problem. Conventional methods either rely on end-to-end RL learning or modular-based policy with topologica…
ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
Tengbo Yu, Guanxing Lu, Zaijia Yang +7
Multi-task robotic bimanual manipulation is becoming increasingly popular as it enables sophisticated tasks that require diverse dual-arm collaboration patterns. Compared to uniman…
Vision Generalist Model: A Survey
Ziyi Wang, Yongming Rao, Shuofeng Sun +8
Recently, we have witnessed the great success of the generalist model in natural language processing. The generalist model is a general framework trained with massive data and is a…