1 paper
Zhichao Xu, Minheng Wang, Yawei Wang +4
Search agents trained with reinforcement learning (RL) interleave reasoning with tool calls in a multi-turn, tool-integrated reasoning (TIR) loop, where each tool invocation return…