9 papers
UniETP: Unifying Environments for Generalizable Embodied Task Planning
Peiran Xu, Jiaqi Zheng, Ziyou Wang +1
This paper focuses on the problem of Embodied Task Planning, where an agent is required to execute a sequence of atomic actions within an interactive environment to complete a user…
GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning
Haozhe Chi, Yang Jin, Yadong Mu
Despite significant progress in agentic long video understanding, existing methods still lack detailed motion comprehension coupled with an efficient memory architecture. In this p…
RePlan-Bot: Multi-Level Replanning for Embodied Instruction Following
Xicheng Gong, Guozheng Sun, Peiran Xu +1
Embodied instruction following (EIF) requires agents to understand and execute complex natural language commands within interactive 3D environments. Despite recent advances, existi…
Extending Embodied Question Answering from Perception to Decision
Xicheng Gong, Qiwei Li, Peiran Xu +1
Embodied Question Answering (EQA) connects perception, reasoning, and interaction within embodied environments. However, existing datasets and benchmarks remain fragmented, each fo…
RotVLA: Rotational Latent Action for Vision-Language-Action Model
Qiwei Li, Xicheng Gong, Xinghang Li +5
Latent Action Models (LAMs) have emerged as an effective paradigm for handling heterogeneous datasets during Vision-Language-Action (VLA) model pretraining, offering a unified acti…
RoboAgent: Chaining Basic Capabilities for Embodied Task Planning
Peiran Xu, Jiaqi Zheng, Yadong Mu
This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accomplish a given task. Although rece…