most citedSCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

Distilled Reinforcement Learning for LLM Post-training

Chen Wang, Zhaochun Li, Jionghao Bai +4

Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL)…

cs.LG20261 cited

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

Chen Wang, Zhaochun Li, Jionghao Bai +3

Reinforcement learning (RL) is a key paradigm for post-training large language models (LLMs), but the widely used Group Relative Policy Optimization (GRPO) often suffers from entro…

cs.LG2026

Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off

Zhaochun Li, Chen Wang, Jionghao Bai +4

The exploration-exploitation (EE) trade-off is a central challenge in reinforcement learning (RL) for large language models (LLMs). With Group Relative Policy Optimization (GRPO),…

cs.LG2025

EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework

Chen Wang, Lai Wei, Yanzhi Zhang +5

Recent advances in reinforcement learning (RL) have significantly enhanced the reasoning capabilities of large language models (LLMs). Group Relative Policy Optimization (GRPO), a…

cs.LG2025

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning

Yanzhi Zhang, Zhaoxi Zhang, Haoxiang Guan +6

Reinforcement learning has emerged as a powerful paradigm for post-training large language models (LLMs) to improve reasoning. Approaches like Reinforcement Learning from Human Fee…