Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
The Easy, the Hard, and the Learnable: Confidence and Difficulty-Adaptive Policy Optimization for LLM Reasoning
Zhanke Zhou, Xiangyu Lu, Chentao Cao +4
RL with verifiable rewards can substantially improve LLM reasoning, yet standard GRPO-style training often treats easy, hard, and learnable questions alike through uniform sampling…
cs.LG2024
LoRA-drop: Efficient LoRA Parameter Pruning based on Output Evaluation
Hongyun Zhou, Xiangyu Lu, Wang Xu +3
Low-Rank Adaptation (LoRA) is currently the most commonly used Parameter-efficient fine-tuning (PEFT) method, it introduces auxiliary parameters for each layer to fine-tune the pre…