6 papers
Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration
Zihan Guo, Zeyi Chen, Zhiyu Chen +15
Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Ther…
Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task
Qianyu Yao, Fei Sun, Bocheng Huang +10
Background. Large language models and AI agents are increasingly used to support biomedical research, but native model outputs may omit key analytical steps, misuse methods, or ove…
SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior
Zhiyu Chen, Zihan Guo, Bo Huang +4
Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a Skill says from how it is organ…
SW--Bench: Benchmarking Autonomous Software Agent Generation for Agentic Web
Linyao Chen, Bo Huang, Qinlao Zhao +13
The Agentic Web is emerging as a paradigm in which autonomous software agents interact with online resources and with each other to accomplish user goals. However, the capacity of…
MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills
Yingyong Hou, Xinyuan Lao, Huimei Wang +10
Background: Agent skills are increasingly deployed as modular, reusable capability units in AI agent systems. Medical research agent skills require safeguards beyond general-purpos…
Holos: A Web-Scale LLM-Based Multi-Agent System for the Agentic Web
Xiaohang Nie, Zihan Guo, Zicai Cui +20
As large language models (LLM)-driven agents transition from isolated task solvers to persistent digital entities, the emergence of the Agentic Web, an ecosystem where heterogeneou…