benchmark 1multimodal large language model 1parallel action planning 1proactive agents 1task graphs 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.RO2026
ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response
Xiaomeng Zhu, Fengming Zhu, Weijie Zhou +8
The paper introduces ProAct-75, a benchmark of 75 proactive tasks with step‑level annotations and task graphs, and presents ProAct-Helper, a multimodal LLM that uses these graphs f…
cs.RO2026
Plan Right, Then Plan Tight: Symbolic RL for Efficient Embodied Reasoning
Xiangli Shi, Xiaomeng Zhu, Ye Tian +5
Embodied task planning asks an agent to turn a natural-language instruction into an executable sequence of actions in a physical scene, and is a building block for household, assis…
cs.AI2025
ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement Learning
Weijie Zhou, Xuangtang Xiong, Ye Tian +8
Multimodal Large Language Models (MLLMs) have empowered embodied agents with remarkable capabilities in planning and reasoning. However, when facing ambiguous natural language inst…