1 paper · 1 filter
Zi-Yi Dou, Cheng-Fu Yang, Xueqing Wu +2
Finetuning language agents with reasoning-action trajectories is effective, but obtaining these trajectories from human annotations or stronger models is costly and sometimes impra…