1 citations · 1 across the 1 of their papers we have counts for
1 paper
Taiwei Shi, Yiyang Wu, Linxin Song +2
Reinforcement finetuning (RFT) has shown great potential for enhancing the mathematical reasoning capabilities of large language models (LLMs), but it is often sample- and compute-…