3 papers
cs.AI2026
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
Zhilin Wang, Han Song, Runzhe Zhan +13
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it…
cs.LG2026
ExGRPO: Learning to Reason from Experience
Runzhe Zhan, Yafu Li, Zhi Wang +5
Reinforcement learning from verifiable rewards (RLVR) is an emerging paradigm for improving the reasoning ability of large language models. However, standard on-policy training dis…
cs.CL2025
Towards an AI Musician: Synthesizing Sheet Music Problems for Musical Reasoning
Zhilin Wang, Zhe Yang, Yun Luo +8
Enhancing the ability of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) to interpret sheet music is a crucial step toward building AI musicians. However,…