103 citations · 118 across the 12 of their papers we have counts for
1 paper · 1 filter
Shyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer +5
RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deployment requires reliable performance across…