6 papers · 1 filter
AdaGamma: State-Dependent Discounting for Temporal Adaptation in Reinforcement Learning
Yaomin Wang, Jianting Pan, Ran Tian +4
The discount factor in reinforcement learning controls both the effective planning horizon and the strength of bootstrapping, yet most deep RL methods use a single fixed value acro…
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
Chaoli Mou, Zhan Zhuang, Xinning Chen +1
Reinforcement Learning with Verifiable Rewards (RLVR) has become a key approach for improving the reasoning abilities of large language models. However, widely used critic-free alg…
One-Token Verification for Reasoning Correctness Estimation
Zhan Zhuang, Xiequn Wang, Zebin Chen +4
Recent breakthroughs in large language models (LLMs) have led to notable successes in complex reasoning tasks, such as mathematical problem solving. A common strategy for improving…
PLAN: Proactive Low-Rank Allocation for Continual Learning
Xiequn Wang, Zhan Zhuang, Yu Zhang
Continual learning (CL) requires models to continuously adapt to new tasks without forgetting past knowledge. In this work, we propose \underline{P}roactive \underline{L}ow-rank \u…
Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation
Zhan Zhuang, Xiequn Wang, Wei Li +9
Low-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal mini…
CopRA: A Progressive LoRA Training Strategy
Zhan Zhuang, Xiequn Wang, Yulong Zhang +3
Low-Rank Adaptation (LoRA) is a parameter-efficient technique for rapidly fine-tuning foundation models. In standard LoRA training dynamics, models tend to quickly converge to a lo…