4 papers · 1 filter
Verifiable Process Rewards for Agentic Reasoning
Huining Yuan, Zelai Xu, Huaijie Wang +6
Reinforcement learning from verifiable rewards (RLVR) has improved the reasoning abilities of large language models (LLMs), but most existing approaches rely on sparse outcome-leve…
MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation
Lu Yang, Zelai Xu, Minyang Xie +4
Large Language Model (LLM) agents have demonstrated remarkable proficiency in learned tasks, yet they often struggle to adapt to non-stationary environments with feedback. While In…
A Survey on Self-play Methods in Reinforcement Learning
Ruize Zhang, Zelai Xu, Chengdong Ma +8
Self-play, a learning paradigm where agents iteratively refine their policies by interacting with historical or concurrent versions of themselves or other evolving agents, has show…
ICPL: Few-shot In-context Preference Learning via LLMs
Chao Yu, Qixin Tan, Hong Lu +5
Preference-based reinforcement learning is an effective way to handle tasks where rewards are hard to specify but can be exceedingly inefficient as preference learning is often tab…