12 citations · 33 across the 21 of their papers we have counts for
6 papers · 1 filter
JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
Yihan Lin, Jiawei He, Shifeng Bao +6
Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAM…
DA-Nav: Direction-Aware City-Scale Vision-Language Navigation
Ye Yuan, Kehan Chen, Xinqiang Yu +7
City-scale outdoor navigation is currently hindered by the heavy reliance on dense maps or costly navigation supervision. In this work, we introduce a novel paradigm for leveraging…
FloorPlan-VLN: A New Paradigm for Floor Plan Guided Vision-Language Navigation
Kehan Chen, Yan Huang, Dong An +5
Existing Vision-Language Navigation (VLN) task requires agents to follow verbose instructions, ignoring some potentially useful global spatial priors, limiting their capability to…
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics
Enshen Zhou, Yibo Li, Jingkun An +12
Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded reasoning compounded with complex spa…
MarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environments
Xu Hu, Yiyang Feng, Junran Peng +7
The development of embodied agents for complex commercial environments is hindered by a critical gap in existing robotics datasets and benchmarks, which primarily focus on househol…
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
Zekun Qi, Wenyao Zhang, Yufei Ding +15
While spatial reasoning has made progress in object localization relationships, it often overlooks object orientation-a key factor in 6-DoF fine-grained manipulation. Traditional p…