3 papers
cs.LG2026
Information-Theoretic Limits of Safety Verification for Self-Improving Systems
Arsenios Scrivens
Can a safety gate permit unbounded beneficial self-modification while maintaining bounded cumulative risk? We formalize this question through dual conditions -- requiring sum delta…
cs.LG2026
Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates
Arsenios Scrivens
Can classifier-based safety gates maintain reliable oversight as AI systems improve over hundreds of iterations? We provide comprehensive empirical evidence that they cannot. On a…
stat.ML2026
Holographic Invariant Storage: Design-Time Safety Contracts via Vector Symbolic Architectures
Arsenios Scrivens
We introduce Holographic Invariant Storage (HIS), a protocol that assembles known properties of bipolar Vector Symbolic Architectures into a design-time safety contract for LLM con…