1 paper
Cong Pang, Xuyu Feng, Yujie Yi +8
Long-horizon reinforcement learning for information seeking agents remains difficult because terminal rewards reveal whether the final answer is correct, but not which acquired inf…