1 paper
Guangya Wan, Zixin Stephen Xu, Sasa Zorc +4
Sampling multiple responses is a common way to improve LLM output quality, but it comes at the cost of additional computation. The key challenge is deciding when to stop generating…