4 papers · 1 filter
IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents
Senhao Wang, Chenghao Cai, Haitao Hu +4
Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as mathematical reasoning and c…
RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning
Yiteng Mao, Kenan Xu, Yijia Lyu +3
While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processe…
AgentSchool: An LLM-Powered Multi-Agent Simulation for Education
Yulei Ye, Wenhao Li, Zhong Wen +23
Despite the rapid deployment of LLMs into classrooms, validating educational AI remains uniquely intractable: interventions act on developing learners whose cognitive and social tr…
See, Plan, Snap: Evaluating Multimodal GUI Agents in Scratch
Xingyi Zhang, Yulei Ye, Kaifeng Huang +2
Block-based programming environments such as Scratch play a central role in low-code education, yet evaluating the capabilities of AI agents to construct programs through Graphical…