7 papers · 1 filter
PATH-Bench: Path-Dependent Evaluation of Lifelong Agents
Xidong Yang, Xingyi Zhang, Wenhao Li +7
Lifelong LLM agents increasingly adapt through external learning states that store past interactions as retrievable memories or reusable skills, yet existing benchmarks rarely acco…
Agentic Episodic Control
Xidong Yang, Wenhao Li, Junjie Sheng +4
Reinforcement learning (RL) remains fundamentally limited by poor data efficiency and weak generalization. Prior episodic RL methods attempt to alleviate this via external memory m…
RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning
Yiteng Mao, Kenan Xu, Yijia Lyu +3
While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processe…
AgentSchool: An LLM-Powered Multi-Agent Simulation for Education
Yulei Ye, Wenhao Li, Zhong Wen +23
Despite the rapid deployment of LLMs into classrooms, validating educational AI remains uniquely intractable: interventions act on developing learners whose cognitive and social tr…
OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation
Haochen Yang, Ke Zhao, Mengyuan Ma +3
Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient paradigm for automated optimiza…
Where Paths Collide: A Comprehensive Survey of Classic and Learning-Based Multi-Agent Pathfinding
Shiyue Wang, Haozheng Xu, Yuhan Zhang +4
Multi-Agent Path Finding (MAPF) is a fundamental problem in artificial intelligence and robotics, requiring the computation of collision-free paths for multiple agents navigating f…