17 papers
Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data
Yangtao Chen, Zixuan Chen, Peiyang Wang +4
Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and scene/embodiment alignment: t…
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
Hongyu Ding, Sizhuo Zhang, Ziming Xu +13
Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environments it has never seen. The dom…
INHerit-SG: Incremental Hierarchical Semantic Scene Graphs with RAG-Style Retrieval
YukTungSamuel Fang, Zhikang Shi, Jiabin Qiu +5
Driven by recent advancements in foundation models, semantic scene graphs have emerged as a promising paradigm for high-level 3D environmental abstraction in robot navigation. Howe…
FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation
Dian Shao, Zhengzheng Xu, Peiyang Wang +4
UAV vision-language navigation (VLN) requires an agent to navigate complex 3D environments from an egocentric perspective while following ambiguous multi-step instructions over lon…
V-Dreamer: Automating Robotic Simulation and Trajectory Synthesis via Video Generation Priors
Songjia He, Zixuan Chen, Hongyu Ding +5
Training generalist robots demands large-scale, diverse manipulation data, yet real-world collection is prohibitively expensive, and existing simulators are often constrained by fi…
ST-VLA: Enabling 4D-Aware Spatiotemporal Understanding for General Robot Manipulation
You Wu, Zixuan Chen, Cunxu Ou +9
Robotic manipulation in open-world environments requires reasoning across semantics, geometry, and long-horizon action dynamics. Existing hierarchical Vision-Language-Action (VLA)…