18 papers
EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking
Jan Kulik, Bjarni Dagur Thor Karason, Yung-Hsu Yang +3
Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured rep…
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
Hongze Wang, Boyang Sun, Jiaxu Xing +5
Object-Goal Navigation (ObjectNav) is a key capability for deploying mobile robots in everyday environments such as homes, schools, and workplaces. In this task, an agent must loca…
LIME: Learning Intent-aware Camera Motion from Egocentric Video
Boyang Sun, Jiajie Li, Yung-Hsu Yang +6
Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region, or obtain a view that responds to a user's intent. While vis…
OpenFrontier: General Navigation with Visual-Language Grounded Frontiers
Esteban Padilla-Cerdio, Boyang Sun, Marc Pollefeys +1
Open-world navigation requires robots to make decisions in complex everyday environments while adapting to flexible task requirements. Conventional navigation approaches often rely…
LocalNav: Distilling Frontier VLMs and Embodied RL for On-Device Object Goal Navigation
Nicolas Baumann, Liam Boyle, Pu Deng +5
Vision Language Models (VLMs) have emerged in the robotic domain as a powerful tool that enables environmental perception with language context, serving as a catalyst for open-voca…
PROSE: Training-Free Egocentric Scene Registration with Vision-Language Models
Zhiang Chen, Nahyuk Lee, Boyang Sun +4
Registering two captures of the same indoor space taken at different times underpins persistent spatial memory for robots and AR systems, yet the realistic version of this task is…