3 papers
cs.LG2026
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2
Hybrid post-training usually combines supervised fine-tuning and reinforcement learning, but fixed mixing schedules cannot adapt when the relative noise of the two signals changes…
stat.ML2026
Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning
Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2
We formalize Rollout Informativeness under a Fixed Budget (RIFB) as the expected non-vanishing policy-gradient mass that a tool-use rollout set injects into Group Relative Policy O…
cs.AI2026
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2
Agent-repair leaderboards reorder under evaluator reconfiguration, and a measurable share of the reordering is produced by methods that consult evaluator-derived signal during inte…