2 papers
cs.LG2026
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
Haokun Liu, Gyung Hyun Je, Marco Ciccone +3
The widespread availability of fine-tuned LoRA modules for open pre-trained models has led to an interest in methods that can adaptively merge LoRAs to improve performance. These m…
cs.LG2026
A Gradient Perspective on RLVR Stability and Winner Advantage Policy Optimization
Prasanth YSS, Zhichen Ren, Rasa Hosseinzadeh +6
Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but GRPO-style optimization remains prone to collapse. We analyse this instability through…