19 citations · 91 across the 20 of their papers we have counts for
1 paper · 2 filters
Jiamian Wang, Ruiyi Zhang, Tong Yu +5
Recent methods train search agents via reinforcement learning from (question, answer, evidence) tuples without requiring expert trajectories. The tuples serve as the training envir…