4 papers
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
Xiaoyuan Liu, Jianhong Tu, Yuqi Chen +26
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, cr…
ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation
Boqin Yuan, Yue Su, Renchu Song +2
Skill-distillation pipelines learn reusable rules from LLM agent trajectories, but they lack a key signal: how much each step costs. Without per-step cost, a pipeline cannot distin…
VULSOLVER: Vulnerability Detection via LLM-Driven Constraint Solving
Xiang Li, Yueci Su, Jiahao Liu +4
Traditional vulnerability detection methods rely heavily on predefined rule matching, which often fails to capture vulnerabilities accurately. With the rise of large language model…
SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents
Kunlun Zhu, Jiaxun Zhang, Ziheng Qi +6
Recent advancements in large language model (LLM) agents have significantly accelerated scientific discovery automation, yet concurrently raised critical ethical and safety concern…