5 papers
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2
Agent-repair leaderboards reorder under evaluator reconfiguration, and a measurable share of the reordering is produced by methods that consult evaluator-derived signal during inte…
Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair
Yuelin Hu, Zhenbo Yu, Zhengxue Cheng +2
Many continual-learning methods modify gradients upstream (e.g., projection, penalty rescaling, replay mixing) while treating Adam as a neutral backend. We show this composition ha…
Benchmarking LLM Tool-Use in the Wild
Peijie Yu, Wei Liu, Yifan Yang +4
Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inherently wild, being intricate,…
WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search
Yuelin Hu, Zhengxue Cheng, Ronghua Wu +5
Specialized web tasks in finance, biomedicine, and pharmaceuticals remain challenging due to missing domain priors: queries drift, evidence is noisy, and reasoning is brittle. We p…
Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models
Yuelin Hu, Zhengxue Cheng, Wei Liu +1
Hybrid training methods for large language models combine supervised fine tuning (SFT) on expert demonstrations with reinforcement learning (RL) on model rollouts, typically at the…