1 paper
Zexu Sun, Yongcheng Zeng, Erxue Min +3
Contemporary progress in large language models (LLMs) has revealed notable inferential capacities via reinforcement learning (RL) employing verifiable reward, facilitating the deve…