1 citations · 1 across the 9 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Distilled Reinforcement Learning for LLM Post-training
Chen Wang, Zhaochun Li, Jionghao Bai +4
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL)…
cs.LG2026★ 1 cited
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
Chen Wang, Zhaochun Li, Jionghao Bai +3
Reinforcement learning (RL) is a key paradigm for post-training large language models (LLMs), but the widely used Group Relative Policy Optimization (GRPO) often suffers from entro…