1 paper
Chenyu Zhou, Qiliang Jiang, Shuning Wu +1
Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattere…