2 papers
cs.LG2026
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Tong Zheng, Skylar Zhai, Zhan Cheng +7
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves rew…
cs.AI2026
Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL
Youling Huang, Tiankuo Xu, Jiaji Liu +10
Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to so…