collaborators

9 papers

cs.LG2026

GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training

Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2

Hybrid post-training usually combines supervised fine-tuning and reinforcement learning, but fixed mixing schedules cannot adapt when the relative noise of the two signals changes…

stat.ML2026

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning

Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2

We formalize Rollout Informativeness under a Fixed Budget (RIFB) as the expected non-vanishing policy-gradient mass that a tool-use rollout set injects into Group Relative Policy O…

cs.AI2026

AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair

Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2

Agent-repair leaderboards reorder under evaluator reconfiguration, and a measurable share of the reordering is produced by methods that consult evaluator-derived signal during inte…

cs.LG2026

Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair

Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2

Many continual-learning methods modify gradients upstream (e.g., projection, penalty rescaling, replay mixing) while treating Adam as a neutral backend. We show this composition ha…

cs.IR2026

WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search

Yuelin Hu, Zhengxue Cheng, Ronghua Wu +5

Specialized web tasks in finance, biomedicine, and pharmaceuticals remain challenging due to missing domain priors: queries drift, evidence is noisy, and reasoning is brittle. We p…

cs.LG2026

Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models

Yuelin Hu, Zhengxue Cheng, Wei Liu +1

Hybrid training methods for large language models combine supervised fine tuning (SFT) on expert demonstrations with reinforcement learning (RL) on model rollouts, typically at the…