1 paper · 1 filter
Shengtian Yang, Yu Li, Shuo He +4
Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a \emph{single} policy network, causing…