1 paper
Chuxuan Hu, Yuxuan Zhu, Antony Kellermann +4
Reinforcement post training (RPT) has recently shown promise in improving the reasoning abilities of large language models (LLMs). However, it remains unclear how well these improv…