2 papers
cs.LG2025
Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Xuerui Su, Shufang Xie, Guoqing Liu +7
Recently, Large Language Models (LLMs) have rapidly evolved, approaching Artificial General Intelligence (AGI) while benefiting from large-scale reinforcement learning to enhance H…
cs.LG2025
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
Xuerui Su, Liya Guo, Yue Wang +4
Inference scaling further accelerates Large Language Models (LLMs) toward Artificial General Intelligence (AGI), with large-scale Reinforcement Learning (RL) to unleash long Chain-…