From the 2 of 71 linked papers with an AI index.
21 papers · 1 filter
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
Xinggang Hu, Chenyangguang Zhang, Alexandros Delitzas +4
The paper introduces a method to build detailed, hierarchical functional 3D scene graphs for indoor environments, using open‑vocabulary visual grounding and temporal graph optimiza…
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
Hongze Wang, Boyang Sun, Jiaxu Xing +5
Object-Goal Navigation (ObjectNav) is a key capability for deploying mobile robots in everyday environments such as homes, schools, and workplaces. In this task, an agent must loca…
LIME: Learning Intent-aware Camera Motion from Egocentric Video
Boyang Sun, Jiajie Li, Yung-Hsu Yang +6
Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region, or obtain a view that responds to a user's intent. While vis…
OpenFrontier: General Navigation with Visual-Language Grounded Frontiers
Esteban Padilla-Cerdio, Boyang Sun, Marc Pollefeys +1
Open-world navigation requires robots to make decisions in complex everyday environments while adapting to flexible task requirements. Conventional navigation approaches often rely…
LocalNav: Distilling Frontier VLMs and Embodied RL for On-Device Object Goal Navigation
Nicolas Baumann, Liam Boyle, Pu Deng +5
Vision Language Models (VLMs) have emerged in the robotic domain as a powerful tool that enables environmental perception with language context, serving as a catalyst for open-voca…
ArtiTwinSplat: Interactable Digital Twin Reconstruction via Gaussian Splatting from RGB-D videos
Pranjal Mishra, René Zurbrügg, Max Wilder-Smith +4
Deploying robots in unstructured real-world environments needs accurate, interactive models of the objects. Constructing these models at scale remains a critical bottleneck for rob…