Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty
Lang Cao, Renhong Chen, Yingtian Zou +9
We introduce the Entropy-Driven Uncertainty Process Reward Model (EDU-PRM), a novel entropy-driven training framework for process reward modeling that enables dynamic, uncertainty-…
cs.LG2026
Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following
Yirong Zeng, Yufei Liu, Xiao Ding +9
A central belief in scaling reinforcement learning with verifiable rewards for instruction following (IF) tasks is that, a diverse mixture of verifiable hard and unverifiable soft…
cs.LG2025
ToolACE: Winning the Points of LLM Function Calling
Weiwen Liu, Xu Huang, Xingshan Zeng +24
Function calling significantly extends the application boundary of large language models, where high-quality and diverse training data is critical for unlocking this capability. Ho…