1 paper
Kevin Lin, Charlie Snell, Yu Wang +4
Scaling test-time compute has emerged as a key ingredient for enabling large language models (LLMs) to solve difficult problems, but comes with high latency and inference cost. We…