3 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Shubham Toshniwal, Ivan Sorokin, Aleksander Ficek +2
Generative reward models with parallel sampling have enabled effective test-time scaling for reasoning tasks. Current approaches employ pointwise scoring of individual solutions or…