1 paper
Pengyu Chen, Shaowei Li, Kai Wang +4
Recent advances in language models have established reinforcement learning as the primary paradigm for eliciting self-correction and long-chain reasoning. While group relative poli…