7 papers
UniETP: Unifying Environments for Generalizable Embodied Task Planning
Peiran Xu, Jiaqi Zheng, Ziyou Wang +1
This paper focuses on the problem of Embodied Task Planning, where an agent is required to execute a sequence of atomic actions within an interactive environment to complete a user…
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi Robotics Team, Jun Guo, Piaopiao Jin +31
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulatio…
RePlan-Bot: Multi-Level Replanning for Embodied Instruction Following
Xicheng Gong, Guozheng Sun, Peiran Xu +1
Embodied instruction following (EIF) requires agents to understand and execute complex natural language commands within interactive 3D environments. Despite recent advances, existi…
Extending Embodied Question Answering from Perception to Decision
Xicheng Gong, Qiwei Li, Peiran Xu +1
Embodied Question Answering (EQA) connects perception, reasoning, and interaction within embodied environments. However, existing datasets and benchmarks remain fragmented, each fo…
RoboAgent: Chaining Basic Capabilities for Embodied Task Planning
Peiran Xu, Jiaqi Zheng, Yadong Mu
This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accomplish a given task. Although rece…
NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
Peiran Xu, Xicheng Gong, Yadong MU
In this work we concentrate on the task of goal-oriented Vision-and-Language Navigation (VLN). Existing methods often make decisions based on historical information, overlooking th…