1 citations · 1 across the 7 of their papers we have counts for
1 paper · 1 filter
Maciej Besta, Leonard Schmidt, Lara Nonino +7
Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards.…