4 papers · 1 filter
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts
Rui Zhang, Xinle Wu, Yao Lu
Reinforcement learning (RL) with verifiable rewards has achieved strong progress in reasoning-oriented LLMs, but extending it to multi-domain RL remains challenging due to reward u…
PowerStep: Memory-Efficient Adaptive Optimization via -Norm Steepest Descent
Yao Lu, Dengdong Fan, Shixun Zhang +1
Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of…
Automatic Configuration of LLM Post-Training Pipelines
Channe Chwa, Xinle Wu, Yao Lu
LLM post-training pipelines that combine supervised fine-tuning and reinforcement learning are difficult to configure under realistic compute budgets: the configuration space is hi…
PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning
Yao Lu, Dengdong Fan, Jianzheng Nie +4
We present PCL-Reasoner-V1.5, a 32-billion-parameter large language model (LLM) for mathematical reasoning. The model is built upon Qwen2.5-32B and refined via supervised fine-tuni…