3 papers
cs.CL2025
Explore Data Left Behind in Reinforcement Learning for Reasoning Language Models
Chenxi Liu, Junjie Liang, Yuqi Jia +4
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an effective approach for improving the reasoning abilities of large language models (LLMs). The Group Relative…
cs.CL2025
On the Convergence of Moral Self-Correction in Large Language Models
Guangliang Liu, Haitao Mao, Bochuan Cao +4
Large Language Models (LLMs) are able to improve their responses when instructed to do so, a capability known as self-correction. When instructions provide only a general and abstr…
cs.CL2024
On the Intrinsic Self-Correction Capability of LLMs: Uncertainty and Latent Concept
Guangliang Liu, Haitao Mao, Bochuan Cao +5
Large Language Models (LLMs) are able to improve their responses when instructed to do so, a capability known as self-correction. When instructions provide only the task's goal wit…