4 papers
PATH-Bench: Path-Dependent Evaluation of Lifelong Agents
Xidong Yang, Xingyi Zhang, Wenhao Li +7
Lifelong LLM agents increasingly adapt through external learning states that store past interactions as retrievable memories or reusable skills, yet existing benchmarks rarely acco…
Reinforced Reasoning for Embodied Planning
Di Wu, Jiaxin Fan, Junzhe Zang +4
Embodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and natural language goals. While recent vision-language models (VLMs)…
TextAtari: 100K Frames Game Playing with Language Agents
Wenhao Li, Wenwu Li, Chuyun Shen +8
We present TextAtari, a benchmark for evaluating language agents on very long-horizon decision-making tasks spanning up to 100,000 steps. By translating the visual state representa…
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines
Pengfei Yu, Dongming Shen, Silin Meng +8
We present RPGBench, the first benchmark designed to evaluate large language models (LLMs) as text-based role-playing game (RPG) engines. RPGBench comprises two core tasks: Game Cr…