3 papers
cs.LG2026
Co-Evolution of Policy and Internal Reward for Language Agents
Xinyu Wang, Hanwei Wu, Jingwei Song +8
Large language model (LLM) agents learn by interacting with environments, but long-horizon training remains fundamentally bottlenecked by sparse and delayed rewards. Existing metho…
cs.LG2025
Incorporating Spatial Information into Goal-Conditioned Hierarchical Reinforcement Learning via Graph Representations
Shuyuan Zhang, Zihan Wang, Xiao-Wen Chang +1
The integration of graphs with Goal-conditioned Hierarchical Reinforcement Learning (GCHRL) has recently gained attention, as intermediate goals (subgoals) can be effectively sampl…
cs.AI2025
SCAR: Shapley Credit Assignment for More Efficient RLHF
Meng Cao, Shuyuan Zhang, Xiao-Wen Chang +1
Reinforcement Learning from Human Feedback (RLHF) is a widely used technique for aligning Large Language Models (LLMs) with human preferences, yet it often suffers from sparse rewa…