1 paper
Qiang Zhang, Ruixue Ding, Fanrui Zhang +9
Reinforcement learning has substantially improved large language model (LLM) agents in verifiable domains, but remains difficult to apply to open-ended agent tasks, where solutions…