1 paper · 1 filter
Xuecheng Liu, Daman Arora, Gokul Swamy +1
While inference-time scaling has improved the reasoning abilities of large language models (LLMs), the need to generate long chains-of-thought (CoTs) is a computational bottleneck.…