collaborators

7 papers

cs.AI2026

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

Zhi Han, Chenxi Zeng, Liuhaichen Yang +3

LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executi…

cs.AI2026

Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration

Zihan Guo, Zeyi Chen, Zhiyu Chen +15

Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Ther…

cs.AI2026

SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior

Zhiyu Chen, Zihan Guo, Bo Huang +4

Agent Skills augment large language model (LLM) agents with procedural knowledge at inference time, but current benchmarks rarely distinguish what a Skill says from how it is organ…

cs.SE2026

Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering

Chenyu Zhou, Huacan Chai, Wenteng Chen +18

Large language model (LLM) agents are increasingly built less by changing model weights than by reorganizing the runtime around them. Capabilities that earlier systems expected the…

cs.CY2026

Synergy: A Next-Generation General-Purpose Agent for Open Agentic Web

Xiaohang Nie, Zihan Guo, Kezhuo Yang +9

AI agents are rapidly expanding in both capability and population: they now write code, operate computers across platforms, manage cloud infrastructure, and make purchasing decisio…

cs.CR2026

SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration

Zihan Guo, Zhiyu Chen, Xiaohang Nie +3

With the rapid evolution of Large Language Model (LLM) agent ecosystems, centralized skill marketplaces have emerged as pivotal infrastructure for augmenting agent capabilities. Ho…