14 papers
Faster-WAM: Do World Action Models Need Deep Action Modules?
Liheng Ma, Rui Heng Yang, Zhanguang Zhang +4
World Action Models (WAMs) couple robot action prediction with video world models. Existing WAMs with shared-backbone and Mixture-of-Transformers designs generally tie the depth of…
RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning
Jinbang Huang, Yuanzhao Hu, Zhiyuan Li +6
Long-horizon robotic tasks require diverse capabilities that no single policy can reliably provide. Heterogeneous policies offer complementary strengths, but orchestrating them req…
Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation
Lingfeng Zhang, Zhanguang Zhang, Liheng Ma +2
End-to-end vision-language navigation (VLN) with causal vision-language models maps instructions and egocentric observations directly to actions, but standard behavior cloning supe…
KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation
Xinyu Shao, Keru Zhou, Guowei Huang +3
Learning manipulation from few demonstrations requires visual priors that capture not only where to interact, but also how the interaction should begin; static priors such as segme…
Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning
Yanzhe Tang, Xinyu Shao, Yuxuan Hu +6
While end-to-end Vision-Language-Action (VLA) models show promise in robotic manipulation, their monolithic paradigm inherently couples semantic reasoning and spatial control. This…
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
Jianzhe Gao, Rui Liu, Yuxuan Xu +6
Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agents commonly encounter percep…