1 paper
Utsav Singh, Sidhaarth Sredharan, Souradip Chakraborty +1
Reinforcement learning with verifiable rewards can improve LLM reasoning, but learning remains sample-inefficient when terminal rewards are sparse. This has motivated a growing lin…