1 paper
Qinhang Wu, Sen Lin, Ming Zhang +2
Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), especially when combined with reinforcement learning (RL) based post-t…