3 papers
cs.AI2026
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution
Zhi Han, Chenxi Zeng, Liuhaichen Yang +3
LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executi…
cs.AI2026
Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration
Zihan Guo, Zeyi Chen, Zhiyu Chen +15
Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Ther…
cs.AI2026
Holos: A Web-Scale LLM-Based Multi-Agent System for the Agentic Web
Xiaohang Nie, Zihan Guo, Zicai Cui +20
As large language models (LLM)-driven agents transition from isolated task solvers to persistent digital entities, the emergence of the Agentic Web, an ecosystem where heterogeneou…