robotics

SoftNav: Injecting 3D Scene Tokens into VLMs for Embodied Navigation

arXiv:2607.14586

summary

SoftNav introduces a method to embed continuous 3D object representations as soft tokens directly into a frozen vision‑language model, enabling efficient embodied navigation with minimal training and strong zero‑shot transfer to new environments.

Abstract

In goal-directed embodied navigation, where an agent must locate a specified target in an unseen environment, 3D scene understanding and navigation reasoning must work in concert. Current approaches transmit 3D scene information to vision-language models (VLMs) through text, suggesting a representation gap in our tested configurations; a controlled ablation confirms that direct embedding-level transfer significantly outperforms the evaluated text serialization formats. We introduce SoftNav, which injects entity-level 3D continuous representations -- one token per detected object or frontier -- into a VLM's hidden space as soft tokens through a lightweight projector. With the 3D encoder and VLM frozen, only ~1,200 samples and ~17M trainable parameters are needed. On HM3D-OVON, SoftNav achieves 74.2%/68.3%/66.7% SR across three splits, surpassing all prior methods in both SR and SPL; the same navigation policy transfers zero-shot to GOAT-Bench (67.2% SR), SG3D (47.2% s-SR), and real-world robot deployment without retraining or architectural modification. Injecting 3D scene tokens directly into VLMs bridges the representation gap, enabling transferable navigation with minimal training.

8 pages, 6 figures. Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026

Topics & keywords

#embodied navigation#vision-language models#3d scene understanding#soft token injection#zero-shot transfersoft tokenscontinuous 3d representationslightweight projectorHM3D-OVONsuccess rate (SR)success weighted by path length (SPL)
SoftNav: Injecting 3D Scene Tokens into VLMs for Embodied Navigation · wovepaper