1 citations · 1 across the 7 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
TreeAdv: Tree-Structured Advantage Redistribution for Group-Based RL
Lang Cao, Hui Ruan, Yongqian Li +5
Reinforcement learning with group-based objectives, such as Group Relative Policy Optimization (GRPO), is a common framework for aligning large language models on complex reasoning…
cs.LG2026
More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty
Lang Cao, Renhong Chen, Yingtian Zou +9
We introduce the Entropy-Driven Uncertainty Process Reward Model (EDU-PRM), a novel entropy-driven training framework for process reward modeling that enables dynamic, uncertainty-…