Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
Xiaoyuan Liu, Jianhong Tu, Yuqi Chen +26
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, cr…
cs.AI2026
ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation
Boqin Yuan, Yue Su, Renchu Song +2
Skill-distillation pipelines learn reusable rules from LLM agent trajectories, but they lack a key signal: how much each step costs. Without per-step cost, a pipeline cannot distin…
cs.AI2025
SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents
Kunlun Zhu, Jiaxun Zhang, Ziheng Qi +6
Recent advancements in large language model (LLM) agents have significantly accelerated scientific discovery automation, yet concurrently raised critical ethical and safety concern…