1 citations · 1 across the 7 of their papers we have counts for
1 paper · 2 filters
Xianlei Zhou, Xiangdi Meng, Yu He +7
Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an action-side Policy-KL regularizer. This puts practitioners in…