3 papers
cs.HC2026
Benchmarking LLM Tool-Use in the Wild
Peijie Yu, Wei Liu, Yifan Yang +4
Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inherently wild, being intricate,…
cs.IR2026
WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search
Yuelin Hu, Zhengxue Cheng, Ronghua Wu +5
Specialized web tasks in finance, biomedicine, and pharmaceuticals remain challenging due to missing domain priors: queries drift, evidence is noisy, and reasoning is brittle. We p…
cs.LG2026
Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models
Yuelin Hu, Zhengxue Cheng, Wei Liu +1
Hybrid training methods for large language models combine supervised fine tuning (SFT) on expert demonstrations with reinforcement learning (RL) on model rollouts, typically at the…