1 citations · 2 across the 6 of their papers we have counts for
1 paper · 1 filter
Soumadeep Saha, Krish Sharma, Akshay Chaturvedi +1
Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significan…