1 paper · 1 filter
Zheyue Tan, Mustapha Abdullahi, Tuo Shi +5
Reinforcement learning (RL) has become a pivotal component of large language model (LLM) post-training, and agentic RL extends this paradigm to operate as agents through multi-turn…