1 paper · 1 filter
Ömer Veysel Çağatan, Barış Akgün, Gözde Gül Şahin +1
Reinforcement learning has become central to post-training large language models, yet dominant algorithms rely on clipping mechanisms that introduce optimization issues at scale, i…