15 papers · 1 filter
LocalNav: Distilling Frontier VLMs and Embodied RL for On-Device Object Goal Navigation
Nicolas Baumann, Liam Boyle, Pu Deng +5
Vision Language Models (VLMs) have emerged in the robotic domain as a powerful tool that enables environmental perception with language context, serving as a catalyst for open-voca…
ArtiTwinSplat: Interactable Digital Twin Reconstruction via Gaussian Splatting from RGB-D videos
Pranjal Mishra, René Zurbrügg, Max Wilder-Smith +4
Deploying robots in unstructured real-world environments needs accurate, interactive models of the objects. Constructing these models at scale remains a critical bottleneck for rob…
Geometric Action Model for Robot Policy Learning
Jisang Han, Seonghu Jeon, Jaewoo Jung +7
Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision-language-acti…
Memory Over Maps: 3D Object Localization Without Reconstruction
Rui Zhou, Xander Yap, Jianwen Cao +3
Target localization is a prerequisite for embodied tasks such as navigation and manipulation. Conventional approaches rely on constructing explicit 3D scene representations to enab…
Articulated 3D Scene Graphs for Open-World Mobile Manipulation
Martin Büchner, Adrian Röfer, Tim Engelbracht +5
Semantics has enabled 3D scene understanding and affordance-driven object interaction. However, robots operating in real-world environments face a critical limitation: they cannot…
Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
Tim Engelbracht, René Zurbrügg, Matteo Wohlrapp +5
We present a dataset for force-grounded, cross-view articulated manipulation that couples what is seen with what is done and what is felt during real human interaction. The dataset…