14 papers
NEO: NeRF It Once, Edit It Many Times for Continuous Object Manipulation
MikoÅaj ZieliÅski, David Hall, Dominik Belter +1
In this paper, we present NEO, a unified framework providing language-guided NeRF editing for robotic manipulation. Our paper introduces (i) a language-guided object removal that c…
TOLiD: Bridging the Architecture Gap in Vision Foundation Model to LiDAR Pretraining via Token Lifting for Distillation
Sutharsan Mahendran, Darshana Priyasad, Kaushik Roy +4
Cross-modal distillation from Vision Foundation Models (VFMs) to LiDAR backbones has recently emerged as a self-supervised pretraining strategy that reduces reliance on dense point…
GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model
Peter Bohm, Saimunur Rahman, Abdelwahed Khamis +3
Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of ho…
Spatially Stratified Distillation for Heterogeneous Radar Place Recognition
Sagun Singh Shrestha, Samuel Harding, Abdelwahed Khamis +2
Scalable, all-weather place recognition increasingly relies on heterogeneous radar place recognition to bridge diverse hardware platforms. A notable application is matching queries…
Visual Place Recognition in Forests with Depth-Aware Distillation
Walter Nedov, Saimunur Rahman, Kavindie Katuwandeniya +3
Visual place recognition in natural forest environments remains challenging due to repetitive vegetation, weak structural cues, and significant appearance variation across traversa…
Cross-Modal Benchmarking for Robotic Perception in Natural Environments
David Hall, Joshua Knights, Mark Cox +1
Natural environments present a complex challenge to robotics perception systems. Current models, particularly vision foundation models, are largely trained on structured, urban env…