paper

Ratchet: How Reliable Must an LLM Judge Be to Retire a Skill?

arXiv:2605.22148

Abstract

A large language model (LLM) agent that writes and edits its own skill library must also decide which skills to keep, from one noisy scalar per skill. The answer is exact: a judge scoring failures as passes at rate or above retires nothing, at any sample size, for eviction margin . Audits find that machinery is rarely built: LLM-written skills are worth percentage points (pp) against a no-skill control, human-written ones pp. Unmaintained, a library enters \emph{library drift}, growing until injecting a skill scores worse than injecting nothing. \textbf{Ratchet} repairs this: it evicts each skill on its measured contribution, caps the library at width , and constrains synthesis, lifting held-out by on a hard MBPP+ slice. The matching non-divergence bound is finite for exactly two reasons, and . Our contribution is the condition this repair carries and no deployed system states. In reference-free domains the scalar comes from an LLM judge, whose two error directions, modelled as a binary channel, behave nothing alike. Passes scored as failures cost sample efficiency, which more trials buy back; failures scored as passes displace the eviction statistic, and no correction inside the rule recovers it. End-task score is a poor alarm, moving by at most a fifth of the governed lift and not monotonically in the rate. We prove both edges of the certifiable region, confirm them in a running loop, and place a judge on a known side in one offline pass.

Code: https://github.com/amazon-science/Self-Evolving-Agents-Ratchet