3d scene understanding 1embodied navigation 1soft token injection 1vision-language models 1zero-shot transfer 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.RO2026
SoftNav: Injecting 3D Scene Tokens into VLMs for Embodied Navigation
Yi Wu, Junjie An, Xiao Liu +6
SoftNav introduces a method to embed continuous 3D object representations as soft tokens directly into a frozen vision‑language model, enabling efficient embodied navigation with m…
cs.RO2024
A Joint Prediction Method of Multi-Agent to Reduce Collision Rate
Mingyi Wang, Hongqun Zou, Yifan Liu +2
Predicting future motions of road participants is an important task for driving autonomously. Most existing models excel at predicting the marginal trajectory of a single agent, bu…