2 papers
cs.CV2026
SR-Nav: Spatial Relationships Matter for Zero-shot Object Goal Navigation
Leyuan Fang, Zan Mao, Zijing Wang +1
Zero-shot object-goal navigation aims to find target objects in unseen environments using only egocentric observation. Recent methods leverage foundation models' comprehension and…
cs.CV2026
PVI: Plug-in Visual Injection for Vision-Language-Action Models
Zezhou Zhang, Songxin Zhang, Xiao Xiong +8
VLA architectures that pair a pretrained VLM with a flow-matching action expert have emerged as a strong paradigm for language-conditioned manipulation. Yet the VLM, optimized for…