1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Hongzhu Yi, Xinming Wang, Zhenghao zhang +12
Within the domain of large language models, reinforcement fine-tuning algorithms necessitate the generation of a complete reasoning trajectory beginning from the input query, which…