2 papers
cs.AI2026
Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning
Bin Lei, Yu Li, Prafulla Kumar Choubey +7
Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differ…
cs.LG2025
Nudging the Boundaries of LLM Reasoning
Justin Chih-Yao Chen, Becky Xiangyu Peng, Prafulla Kumar Choubey +4
Current online reinforcement learning (RL) algorithms like GRPO share a key limitation in LLM reasoning: they cannot learn from problems that are "unsolvable" to the model. In othe…