1 paper · 1 filter
Zibin Meng, Kani Chen
Self-distilled agentic reinforcement learning augments trajectory-level reward with a token-level distillation loss, using as its teacher the same policy conditioned on privileged…