15 citations · 15 across the 3 of their papers we have counts for
1 paper · 1 filter
Shengtian Yang, Yu Li, Shuo He +4
Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a \emph{single} policy network, causing…