activity
20242026
collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

AdaGamma: State-Dependent Discounting for Temporal Adaptation in Reinforcement Learning

Yaomin Wang, Jianting Pan, Ran Tian +4

The discount factor in reinforcement learning controls both the effective planning horizon and the strength of bootstrapping, yet most deep RL methods use a single fixed value acro…

cs.LG2026

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR

Chaoli Mou, Zhan Zhuang, Xinning Chen +1

Reinforcement Learning with Verifiable Rewards (RLVR) has become a key approach for improving the reasoning abilities of large language models. However, widely used critic-free alg…

cs.LG2026

One-Token Verification for Reasoning Correctness Estimation

Zhan Zhuang, Xiequn Wang, Zebin Chen +4

Recent breakthroughs in large language models (LLMs) have led to notable successes in complex reasoning tasks, such as mathematical problem solving. A common strategy for improving…

cs.LG2025

PLAN: Proactive Low-Rank Allocation for Continual Learning

Xiequn Wang, Zhan Zhuang, Yu Zhang

Continual learning (CL) requires models to continuously adapt to new tasks without forgetting past knowledge. In this work, we propose \underline{P}roactive \underline{L}ow-rank \u…

cs.LG2025

Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation

Zhan Zhuang, Xiequn Wang, Wei Li +9

Low-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal mini…

cs.LG2024

CopRA: A Progressive LoRA Training Strategy

Zhan Zhuang, Xiequn Wang, Yulong Zhang +3

Low-Rank Adaptation (LoRA) is a parameter-efficient technique for rapidly fine-tuning foundation models. In standard LoRA training dynamics, models tend to quickly converge to a lo…