credit assignment 1graph-structured retrieval 1question answering 1reinforcement learning 1search agents 1
From the 1 of 6 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
Chenlu Ye, Zhou Yu, Ziji Zhang +5
Reinforcement Learning with Verifiable Rewards (RLVR) improves final-answer accuracy on reasoning tasks, but it does not reliably improve reasoning quality. Because outcome rewards…
cs.LG2026
Decoupled Travel Planning with Behavior Forest
Duanyang Yuan, Sihang Zhou, Yanning Hou +7
Behavior sequences, composed of executable steps, serve as the operational foundation for multi-constraint planning problems such as travel planning. In such tasks, each planning s…