1 citations · 1 across the 18 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Multi-Rollout On-Policy Distillation via Peer Successes and Failures
Weichen Yu, Xiaomin Li, Yizhou Zhao +8
Large language models are often post-trained with sparse verifier rewards, which indicate whether a sampled trajectory succeeds but provide limited guidance about where reasoning s…
cs.LG2026
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning
Binghang Lu, Zheyuan Deng, Runyu Zhang +6
A central challenge in continual learning for large language models (LLMs) is catastrophic forgetting, where adapting to new tasks can substantially degrade performance on previous…
cs.LG2025
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
Zhiwei Zhang, Hui Liu, Xiaomin Li +10
Reward models trained on human preference data have demonstrated strong effectiveness in aligning Large Language Models (LLMs) with human intent under the framework of Reinforcemen…