3 papers
cs.LG2026
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning
Wenwu Fan, Qihong Lin, Zhijie Xia +4
Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference. This training-inference…
cs.LG2026
Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
Yuhang He, Haodong Wu, Siyi Liu +7
Reinforcement Learning with Verifiable Rewards (RLVR) improves the reasoning ability of Large Language Models (LLMs), but sparse outcome rewards make token-level credit assignment…
cs.LG2026
Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models
Longteng Zhang, Sen Wu, Shuai Hou +7
Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in…