1 paper · 1 filter
Chen Jin, Ryutaro Tanno, Tom Diethe +1
Large Language Models (LLMs) often rely on test-time scaling via parallel decoding (for example, 512 samples) to boost reasoning accuracy, but this incurs substantial compute. We i…