4 papers
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution
Zhi Han, Chenxi Zeng, Liuhaichen Yang +3
LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executi…
Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness
Yang Li, Hai Liu, Dian Shao +8
Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation…
StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems
Sergey Volkov, Yang Li, Ye Luo
Agent systems accumulate conflicting observations across branches, retries, and replicas, yet many practical memory layers still collapse disagreement behind overwrite rules that a…
Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration
Zihan Guo, Zeyi Chen, Zhiyu Chen +15
Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Ther…