1 paper
Shihao Cai, Runnan Fang, Jialong Wu +10
Conducting reinforcement learning (RL) in simulated environments offers a cost-effective and highly scalable way to enhance language-based agents. However, previous work has been l…