1 paper
Muheng Li, Jian Qian, Wenlong Mou
Large language models are increasingly deployed with test-time strategies: sample N responses, score them with a reward model or verifier, and return the best. This deployment ru…