1 paper
Xiaojun Wu, Cehao Yang, Honghao Liu +7
Reinforcement learning (RL) for terminal agents needs executable training environments with reliable rewards and useful difficulty. Fixed recipes such as few-shot, Self-Instruct, a…