1 paper · 1 filter
Kaiming Jin, Yuefan Wu, Shengqiong Wu +3
Vision-and-Language Scene navigation is a fundamental capability for embodied human-AI collaboration, requiring agents to follow natural language instructions to execute coherent a…