Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL
Guanqun Zhao, Zijun Xie, Binbin Zheng +3
Asynchronous reinforcement learning has become the standard way to scale training for large language models (LLM), but the resulting policy lag biases the critic toward the stale b…
cs.LG2026
ACPO: Asymmetric Credit Policy Optimization via Mode-Local Entropy Surrogate
Zijun Xie, Yuyang You, Yongzhi Li +8
Outcome-supervised reinforcement learning scales to verifiable reasoning tasks, but trajectory-level rewards assign the same outcome signal to all sampled tokens, overlooking their…
cs.LG2026
ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL
Zijun Xie, Binbin Zheng, Enlei Gong +7
Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Context-management methods make such rollou…