4 papers
IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents
Senhao Wang, Chenghao Cai, Haitao Hu +4
Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as mathematical reasoning and c…
RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning
Yiteng Mao, Kenan Xu, Yijia Lyu +3
While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processe…
See, Plan, Snap: Evaluating Multimodal GUI Agents in Scratch
Xingyi Zhang, Yulei Ye, Kaifeng Huang +2
Block-based programming environments such as Scratch play a central role in low-code education, yet evaluating the capabilities of AI agents to construct programs through Graphical…
Epitome: Pioneering an Experimental Platform for AI-Social Science Integration
Jingjing Qu, Kejia Hu, Jun Zhu +9
Large Language Models (LLMs) enable unprecedented social science experimentation by creating controlled hybrid human-AI environments. We introduce Epitome (www.epitome-ai.com), an…