21 papers
PATH-Bench: Path-Dependent Evaluation of Lifelong Agents
Xidong Yang, Xingyi Zhang, Wenhao Li +7
Lifelong LLM agents increasingly adapt through external learning states that store past interactions as retrievable memories or reusable skills, yet existing benchmarks rarely acco…
Agentic Episodic Control
Xidong Yang, Wenhao Li, Junjie Sheng +4
Reinforcement learning (RL) remains fundamentally limited by poor data efficiency and weak generalization. Prior episodic RL methods attempt to alleviate this via external memory m…
RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning
Yiteng Mao, Kenan Xu, Yijia Lyu +3
While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processe…
Mean-Field Diffuser: Scaling Offline MARL to Thousands of Agents
Wenhao Li, Xiangfeng Wang, Bo Jin
Diffusion-based planning has achieved strong results in single-agent offline reinforcement learning, yet scaling to many-agent systems remains intractable due to the curse of dimen…
AgentSchool: An LLM-Powered Multi-Agent Simulation for Education
Yulei Ye, Wenhao Li, Zhong Wen +23
Despite the rapid deployment of LLMs into classrooms, validating educational AI remains uniquely intractable: interventions act on developing learners whose cognitive and social tr…
OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation
Haochen Yang, Ke Zhao, Mengyuan Ma +3
Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient paradigm for automated optimiza…