1 paper · 1 filter
Zhizheng Jiang, Kang Zhao, Weikai Xu +5
Large reasoning models (LRMs) aim to solve diverse and complex problems through structured reasoning. Recent advances in group-based policy optimization methods have shown promise…