1 paper
Chengzhuo Tong, Ziyu Guo, Renrui Zhang +5
Recent advancements underscore the significant role of Reinforcement Learning (RL) in enhancing the Chain-of-Thought (CoT) reasoning capabilities of large language models (LLMs). T…