7 papers
ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures
Haoran Xu, Lechao Zhang, Daoguo Dong +2
Constructing simulation-ready 3D scenes from multi-view captures is a key bottleneck for Embodied Artificial Intelligence, as downstream tasks require object-level structure, expli…
Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation
Yang Gao, Wuyang Li, Po-Chien Luan +1
Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human-centric, nonrigid agents. However, existing weakly supervise…
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
Hongyu Ding, Sizhuo Zhang, Ziming Xu +13
Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environments it has never seen. The dom…
INHerit-SG: Incremental Hierarchical Semantic Scene Graphs with RAG-Style Retrieval
YukTungSamuel Fang, Zhikang Shi, Jiabin Qiu +5
Driven by recent advancements in foundation models, semantic scene graphs have emerged as a promising paradigm for high-level 3D environmental abstraction in robot navigation. Howe…
LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
Hongyu Ding, Ziming Xu, Yudong Fang +6
LaViRA: Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an agent to navigate unseen environments based on natural language instructions withou…
SEA: Semantic Map Prediction for Active Exploration of Uncertain Areas
Hongyu Ding, Xinyue Liang, Yudong Fang +7
In this paper, we propose SEA, a novel approach for active robot exploration through semantic map prediction and a reinforcement learning-based hierarchical exploration policy. Unl…