Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
Jiashu Yao, Heyan Huang, Daiqing Wu +2
To encourage diverse exploration in reinforcement learning (RL) for large language models (LLMs) without compromising accuracy, we propose Policy Split, a novel paradigm that bifur…
cs.CL2025
Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
Jiashu Yao, Heyan Huang, Shuang Zeng +6
Through reinforcement learning (RL) with outcome correctness rewards, large reasoning models (LRMs) with scaled inference computation have demonstrated substantial success on compl…