1 paper
Xinji Mai, Haotian Xu, Zhong-Zhi Li +5
Large Language Models (LLMs) often struggle with mathematical reasoning tasks requiring precise, verifiable computation. While Reinforcement Learning (RL) from outcome-based reward…