From the 1 of 34 linked papers with an AI index.
4 papers · 1 filter
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning
Yifu Huo, Shunjie Xing, Chenglong Wang +8
Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge is credit assignment, which aim…
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction
Chenglong Wang, Ziming Zhu, Yifu Huo +9
Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranki…
FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models
Kaiyang Ye, Yuan Ge, Junxiang Zhang +8
While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored.…
Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence
Xinyu Liu, Kechen Jiao, Chunyang Xiao +10
On-policy distillation (OPD) has become a promising paradigm for reasoning-oriented post-training of large language models (LLMs), especially when combined with reinforcement learn…