2 papers
cs.CY2026
Agentic Safety is an Epistemic Property, Not a Behavioral One
Charles L. Wang, Keir Dorchen, Peter Jin
Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods are necessary, but they prima…
cs.AI2026
On The Statistical Limits of Self-Improving Agents
Charles L. Wang, Keir Dorchen, Peter Jin
We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework, we prove a sharp boundary: und…