2 papers
cs.CR2026
Basic Legibility Protocols Improve Trusted Monitoring
Ashwin Sreevatsa, Sebastian Prasanna, Cody Rushing
The AI Control research agenda aims to develop control protocols: safety techniques that prevent untrusted AI systems from taking harmful actions during deployment. Because human o…
cs.CR2025
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
Yeonwoo Jang, Shariqah Hossain, Ashwin Sreevatsa +1
In this work, we demonstrate that certain machine unlearning methods may fail under straightforward prompt attacks. We systematically evaluate eight unlearning techniques across th…