4 citations · 11 across the 10 of their papers we have counts for
1 paper · 1 filter
Guoheng Sun, Ziyao Wang, Bowei Tian +7
As post-training techniques evolve, large language models (LLMs) are increasingly augmented with structured multi-step reasoning abilities, often optimized through reinforcement le…