1 paper · 1 filter
Yanhao Li, Lu Ma, Jiaran Zhang +3
Existing approaches typically rely on fixed length penalties, but such penalties are hard to tune and fail to adapt to the evolving reasoning abilities of LLMs, leading to suboptim…