Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Information-Theoretic Limits of Safety Verification for Self-Improving Systems
Arsenios Scrivens
Can a safety gate permit unbounded beneficial self-modification while maintaining bounded cumulative risk? We formalize this question through dual conditions -- requiring sum delta…
cs.LG2026
Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates
Arsenios Scrivens
Can classifier-based safety gates maintain reliable oversight as AI systems improve over hundreds of iterations? We provide comprehensive empirical evidence that they cannot. On a…