2 citations · 2 across the 10 of their papers we have counts for
1 paper · 2 filters
Jinluan Yang, Yuxin Liu, Zhengyu Chen +7
Training tool-use agents typically relies on outcome-based filtering: Supervised Fine-Tuning (SFT) on successful trajectories and Reinforcement Learning (RL) on pass-rate-selected…