1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Xiao Ma, Zhiquan Hu, Yi Wei +6
Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups provide no comparative signal and may hi…