1 paper
Linas Nasvytis, Simon Jerome Han, Ben Prystawski +3
Language models can use verifiable rewards to improve at a wide variety of reasoning tasks. However, both parametric (e.g. RLVR) and non-parametric (e.g. prompt optimization) appro…