2 citations · 4 across the 11 of their papers we have counts for
1 paper · 1 filter
Yang Tian, Rui Wang, Xumeng Wen +5
Long-horizon agentic tasks pose a fundamental credit assignment challenge for outcome-base reinforcement learning: trajectory-level rewards verify final correctness but provide lim…