5 citations · 14 across the 20 of their papers we have counts for
4 papers · 1 filter
From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
Xitai Jiang, Zihan Tang, Wenze Lin +3
Reinforcement learning from verifiable rewards (RLVR) has shown strong promise for LLM reasoning, but outcome-based RLVR remains inefficient on hard problems because correct final-…
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Andrew Zhao, Yiran Wu, Yang Yue +7
Reinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning capabilities of large language models by learning directly from outcome-based rew…
Towards Understanding the Benefit of Multitask Representation Learning in Decision Process
Rui Lu, Yang Yue, Andrew Zhao +2
Multitask Representation Learning (MRL) has emerged as a prevalent technique to improve sample efficiency in Reinforcement Learning (RL). Empirical studies have found that training…
Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL
Yang Yue, Rui Lu, Bingyi Kang +2
The divergence of the Q-value estimation has been a prominent issue in offline RL, where the agent has no access to real dynamics. Traditional beliefs attribute this instability to…