1 paper
Yan Chen, Long Li, Teng Xi +2
Reinforcement learning (RL) has proven highly effective in eliciting the reasoning capabilities of large language models (LLMs). Inspired by this success, recent studies have explo…