1 paper
Wenshuo Zhao, Qi Zhu, Xingshan Zeng +4
An effective way to scale up test-time compute of large language models is to sample multiple responses and then select the best one, as in Grok Heavy and Gemini Deep Think. Existi…