collaborators

12 papers

cs.AI2026

Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

Can Wang, Haoran Chen, Li Yu +4

The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role as the primary interfa…

cs.AI2026

From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

Can Wang, Haoran Chen, Haowen Gao +3

Deep research benchmarks require expert-level tasks and reliable evaluation grounded in task-specific knowledge. Existing benchmarks rely heavily on expert authoring or pre-existin…

cs.AI2026

LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

Anqi Zou, Han Deng, Chengyu Zhang +9

Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over co…

cs.LG2026

A Gradient Perspective on RLVR Stability and Winner Advantage Policy Optimization

Prasanth YSS, Zhichen Ren, Rasa Hosseinzadeh +6

Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but GRPO-style optimization remains prone to collapse. We analyse this instability through…

cs.AI2026

SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing

Haowen Gao, Haoran Chen, Can Wang +5

Agent skills are structured procedural packages that guide frozen LLM agents in specialized workflows. Skills rarely remain sufficient after deployment: edge cases, API changes, an…

cs.AI2026

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning

Qingxu Fu, Boyin Liu, Shuchang Tao +5

Training reinforcement learning (RL) policies for large language model (LLM) agents requires optimizing multi-turn trajectories that interact with external environments. Existing t…