1 paper
Yantao Liu, Zijun Yao, Rui Min +3
Best-of-N (BoN) sampling, a common strategy for test-time scaling of Large Language Models (LLMs), relies on reward models to select the best candidate solution from multiple gener…