6 papers
GUIDE: Goal-Initialized Directional Understanding for End-to-End Visual Navigation
Liang Wang, Jin Jin, KanZhong Yao +6
Learning-based visual navigation for legged robots typically relies on continuous goal updates from hierarchical state estimation to provide a persistent directional reference. Thi…
When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation Detection
Tao Yu, Yujia Yang, Shenghua Chai +17
Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spliced across sources, or augme…
EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation
Yuan Xu, Jiabing Yang, Xiaofeng Wang +16
Imitation learning based visuomotor policies have achieved strong performance in robotic manipulation, yet they often remain sensitive to egocentric viewpoint shifts. Unlike third-…
FloorPlan-VLN: A New Paradigm for Floor Plan Guided Vision-Language Navigation
Kehan Chen, Yan Huang, Dong An +5
Existing Vision-Language Navigation (VLN) task requires agents to follow verbose instructions, ignoring some potentially useful global spatial priors, limiting their capability to…
BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
Yixiang Chen, Peiyan Li, Jiabing Yang +8
Embodied world models have emerged as a promising paradigm in robotics, most of which leverage large-scale Internet videos or pretrained video generation models to enrich visual an…
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
Kehan Chen, Dong An, Yan Huang +5
We address the task of Vision-Language Navigation in Continuous Environments (VLN-CE) under the zero-shot setting. Zero-shot VLN-CE is particularly challenging due to the absence o…