1 paper
Jiaqi Shao, Yuxiang Lin, Munish Prasad Lohani +2
Recent work has explored training Large Language Model (LLM) search agents with reinforcement learning (RL) for open-domain question answering (QA). However, most evaluations focus…