2 papers
cs.AI2026
EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning
Weiyuan Li, Aili Chen, Xintao Wang +11
Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop:…
cs.CL2026
SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design
Jianing Wang, Xintao Wang, Aili Chen +7
Social intelligence enables agents to read social context, infer intent, and adapt over sustained dialogue. As language models become autonomous collaborators, it is central to bui…