works on

From the 1 of 12 linked papers with an AI index.

collaborators
Showing cs.AIShow all

8 papers · 1 filter

cs.AI2026

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

Xing Zhang, Guanghui Wang, Yanwei Cui +4

The paper introduces a framework that co‑evolves evaluation metrics and the skills of LLM agents using an evolutionary loop guided by anchored reference sets, enabling transparent…

cs.AI2026

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

Xing Zhang, Yanwei Cui, Guanghui Wang +4

A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps…

cs.AI2026

Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

Xing Zhang, Yanwei Cui, Guanghui Wang +4

Self-evolving skill libraries face a silent failure mode we term \emph{library drift}: unbounded skill accumulation without outcome-driven lifecycle management causes retrieval deg…

cs.AI2026

Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents

Xing Zhang, Guanghui Wang, Yanwei Cui +4

As LLM agents scale to long-horizon, multi-session deployments, efficiently managing accumulated experience becomes a critical bottleneck. Agent memory systems and agent skill disc…

cs.AI2026

Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement Learning

Yanwei Cui, Xing Zhang, Yulong Zhang +4

Training-free verbal reinforcement learning enables LLM agents to learn from world feedback -- objective signals such as dynamic task outcomes, market returns, or demand forecasts…

cs.AI2026

Guardrails Beat Guidance: A Large-Scale Study of Rules, Skills, and Persistent Configuration for Coding Agents

Xing Zhang, Guanghui Wang, Yanwei Cui +4

Random rules improve a coding agent's task performance as much as expert-curated ones (both pp on a discriminative subset of SWE-bench Verified), and in our data every indiv…