15 papers
Faster-WAM: Do World Action Models Need Deep Action Modules?
Liheng Ma, Rui Heng Yang, Zhanguang Zhang +4
World Action Models (WAMs) couple robot action prediction with video world models. Existing WAMs with shared-backbone and Mixture-of-Transformers designs generally tie the depth of…
RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning
Jinbang Huang, Yuanzhao Hu, Zhiyuan Li +6
Long-horizon robotic tasks require diverse capabilities that no single policy can reliably provide. Heterogeneous policies offer complementary strengths, but orchestrating them req…
Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation
Lingfeng Zhang, Zhanguang Zhang, Liheng Ma +2
End-to-end vision-language navigation (VLN) with causal vision-language models maps instructions and egocentric observations directly to actions, but standard behavior cloning supe…
Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation
Jinbang Huang, Zhiyuan Li, Yuanzhao Hu +4
Large Language Models (LLMs) have recently shown strong promise for robotic task planning, particularly through automatic planning domain generation. However, prior approaches larg…
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
Jianzhe Gao, Rui Liu, Yuxuan Xu +6
Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agents commonly encounter percep…
Do World Action Models Generalize Better than VLAs? A Robustness Study
Zhanguang Zhang, Zhiyuan Li, Behnam Rahmati +11
Robot action planning in the real world is challenging as it requires not only understanding the current state of the environment but also predicting how it will evolve in response…