3 papers
cs.CL2026
Agent Skill Evaluation and Evolution: Frameworks and Benchmarks
Kexin Ding, Yang Zhou, Can Jin +3
The growth of agent skills has transformed how agentic systems are built, evaluated, and deployed. As skill libraries continue to scale, rigorous evaluation becomes critical to ens…
cs.AI2026
Evidence Over Plans: Online Trajectory Verification for Skill Distillation
Yang Zhou, Zihan Dong, Zhenting Wang +7
Agent skills can remarkably improve task success rates by using human-written procedural documents, but their quality is difficult to assess without environment-grounded verificati…
cs.LG2026
Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation
Bruce Changlong Xu, Adarsh Kumarappan, Mu Zhou
Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy w…