1 paper
Gijs Kassenaar, Zhao Yang, Vincent François-Lavet
Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computatio…