8 papers
Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Model
Zizhao Yuan, Zhengtu Liang, Taowen Wang +7
Recent advances in action-conditioned world models show promising progress in modeling complex interactions and forecasting future states under diverse action sequences. While thes…
HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning
Taowen Wang, Zikang Xie, Bin Yang +13
Humanoid robots promise whole-body interaction in human-centered environments, but scalable policy learning remains difficult because task-level decision-making and whole-body dyna…
CoRe-MoE: Contrastive Reweighted Mixture of Experts for Multi-Terrain Humanoid Locomotion with Gait Adaptation
Kailun Huang, Zikang Xie, Yanzhe Xie +8
Humans primarily rely on walking and running to traverse complex terrains. Similarly, humanoid robots should be able to smoothly transition between walking and running while mainta…
What Limits Vision-and-Language Navigation ?
Yunheng Wang, Yuetong Fang, Taowen Wang +9
Vision-and-Language Navigation (VLN) is a cornerstone of embodied intelligence. However, current agents often suffer from significant performance degradation when transitioning fro…
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
Jiaxi Zhang, Yunheng Wang, Wei Lu +8
3D Visual Grounding (3DVG) localizes objects from natural-language descriptions in 3D scenes and is fundamental to embodied AI applications. Although foundation models enable open-…
HERO: Hierarchical Traversable 3D Scene Graphs for Embodied Navigation Among Movable Obstacles
Yunheng Wang, Yixiao Feng, Yuetong Fang +5
3D Scene Graphs (3DSGs) constitute a powerful representation of the physical world, distinguished by their abilities to explicitly model the complex spatial, semantic, and function…