1 paper
Aakash Lahoti, Kevin Y. Li, Berlin Chen +5
Scaling inference-time compute has emerged as an important driver of LLM performance, making inference efficiency a central focus of model design alongside model quality. While the…